The M3 team is interested in models and methods aimed at the fundamental description and automatic processing of natural language and its multilingual dimension. This involves the development of models and methods able to process, extract, and describe relevant characteristics of these systems from datasets. We consider languages in all their uses, with special focus on under-resourced languages and comparative approaches to language description (multilingual approaches). The team questions the interactions between models and systems: How does a model help our understanding of languages? How do sets of languages differ, and at what levels? How do we use generative models to produce controlled instances of given phenomena? The implementation of explainable and adapted models allows interaction between disciplines and promotes interdisciplinary work (computer science, language sciences, sociology, psychology). The research carried out in the M3 team falls into three main axes:
A- Models and Methods
This axis focuses on learning paradigms: developing data models and algorithms (models with many or few parameters, generative or not), with a view to their application to automatic language processing and languages as structured objects. These models are typically applied to the objects studied in the other two axes. Particular attention is paid to issues related to accessibility: reflecting on the specific methods to be implemented to develop inclusive technologies for our society. Effective, sober models, adapted to the representation of specific data, are particularly sought to promote explainable and responsible approaches to the data studied. These models make it possible to create hybrid solutions that try to control generative AIArtificial Intelligence. They also provide approaches that can be used to build more ethical systems.
B- Typology, Variation, and Universals in Languages
This axis focuses on describing the characteristics of linguistic systems based on corpora. This involves automatically applying typological schemes and language comparisons according to these characteristics. Considering variation is a major point, with work on diatopic, diastratic, diaphasic, or diachronic changes within languages, or linked to language contact of under-resourced or well-described languages. Syntactic, phonological, phonetic, articulatory, and prosodic systems are considered. The representation of the differences (in terms of distances or projected onto an atlas) between the systems studied is another highlight.
C- Contextualized Behaviors for Interaction
This axis models and describes performances during situated communicational interactions, at para- and extra-linguistic levels: whether for pragmatic functions (speech acts, attitudinal nuances), emotions (affective interaction, social emotions), nudges (gentle manipulation), vocal effort (Lombard speech, voice strength), etc. It aims to propose models of behavioral changes linked to these phenomena, in order to be able to detect them, measure their variation or dynamics, and categorize them. The analysis of the acoustic-linguistic parameters of the voice (parameters derived from models, glottal source, articulatory choices, etc.) makes it possible to link performances and functions.dèles, source glottique, choix articulatoires, etc.) permet de faire le lien entre les performances et les fonctions.
L’équipe se compose de 9 membres permanents (chercheurs CNRS, enseignants-chercheurs à l’Université Paris-Saclay), 6 doctorants, et 9 personnes ingénieures ou chercheurs CDD. Nous entretenons des liens avec les industriels (thèses en contrat CIFRE, projets de recherche) et organisons régulièrement des manifestations scientifiques (conférence TALN, ateliers et workshops scientifiques, etc.).
Younes Boufouss, Luc Pommeret. Natural Language Inference using Enhanced Knowledge Graphs for Multi-Hop Reasoning. JDSE 2026 – 11th Junior Conference on Data Sciences and Engineering, Sep 2026, Orsay, France. . ⟨hal-05765575⟩
Lucía Catalán Gris, Kim Gerdes, John S Y Lee. The Linguist’s Lie Detector: Linguistic Knowledge in Large Language Models. NSLP 2026 @ LREC 2026 – 3rd International Workshop on Natural Scientific Language Processing, May 2026, Palma de Mallorca, Spain. pp.235-246, ⟨10.63317/2f7zobe6yo7h⟩. ⟨hal-05652276⟩
Jianying Liu, Kim Gerdes, Jean-Marc Deltorn. Beyond frequency measures: Can contextual embeddings capture meaning change in scientific texts?. JADT 2026 – 18th International Conference on Statistical Analysis of Textual Data, Jul 2026, Palermo, Italy. pp.232-241. ⟨hal-05731086⟩
Marta López-Rauhut, Loic Landrieu, Mathieu Aubry, Anne-Laure Ligozat. Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model. Transactions on Machine Learning Research , 2026. ⟨hal-05749590⟩
Younes Boufouss, Luc Pommeret, Thomas Gerald, Patrick Paroubek, Sophie Rosset. Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?. AKBC @ EMNLP, ACL, Oct 2026, Budapest, Hungary. ⟨hal-05745292v3⟩
Anisia Popescu, Lori Lamel, Marc Evrard, Ioana Vasilescu. Tracking /r/ Deletion: Forced Alignment of Pronunciation Variants and Sociophonetic Insights into Post-Obstruent Final /r/ in French. Interspeech 2025, Aug 2025, Rotterdam, Netherlands. pp.2945 – 2949, ⟨10.21437/interspeech.2025-967⟩. ⟨hal-05682650⟩
Agnieszka Dryjańska, Till Überrück-Fries, Agata Savary. Identification and annotation of multiword expressions in an end-user application: The case of teaching/ learning French as a foreign language. Verginica Barbu Mititelu; Voula Giouli. Multiword expressions in Natural Language Processing: Current trends and challenges, 8, Language Science Press, pp.225-254, 2026, Phraseology and Multiword Expressions, 978-3-96110-591-5. ⟨10.5281/zenodo.21277679⟩. ⟨hal-05734416⟩
Dylan Sechet, Marc Evrard, Matthieu Kowalski. Simulation-Based Inference for Plate Reverb System Identification. DAFX 2026 – 29th International Conference on Digital Audio Effects, Sep 2026, Cambridge, United States. ⟨hal-05742718⟩
Alex Colagrande, Paul Caillon, Eva Feillet, Alexandre Allauzen. Linear Attention with Global Context: A Multipole Attention Mechanism for Vision and Physics. ICCVW 2025 – IEEE/CVF International Conference on Computer Vision Workshops, Oct 2025, Honolulu, United States. pp.3130-3139, ⟨10.1109/ICCVW69036.2025.00326⟩. ⟨hal-05742260⟩
Alex Colagrande, Paul Caillon, Eva Feillet, Alexandre Allauzen. Limits of Resolution Equivariance in Fourier Neural Operators. AIArtificial Intelligence&PDE: ICLR 2026 Workshop on AIArtificial Intelligence and Partial Differential Equations, Apr 2026, Rio de Janeiro, Brazil. arXiv, 2026, ⟨10.48550/arXiv.2606.00677⟩. ⟨hal-05742295⟩
Adèle Jatteau, Jinyu Li, Lori Lamel. Y a-t-il des géminées en français ? Une étude acoustique des mots en dans de grands corpus de parole. Congrès mondial de linguistique française, Jul 2026, Arras, France. pp.09008, ⟨10.1051/shsconf/202623209008⟩. ⟨hal-05740510⟩
Ayoub Hammal, Pierre Zweigenbaum, Caio Corro. On the Rejection Criterion for Proxy-based Test-time Alignment. ACL 2026 – 64th Annual Meeting of the Association for Computational Linguistics, Jul 2026, San Diego, United States. pp.547-554, ⟨10.18653/v1/2026.acl-short.46⟩. ⟨hal-05689863⟩
Clémentine Bleuze, Karën Fort, Vincent P Martin, Aurélie Névéol. Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024. JMIR AIArtificial Intelligence, 2026, 5, pp.e88082-e88082. ⟨10.2196/88082⟩. ⟨hal-05717650⟩
Marco Naguib, Christel Gérardin, Victor Beaucoté, Cyril Charron, Adrien Joseph, et al.. Evaluating the Retrieval Component in a Retrieval-Augmented Summarization System for Patient Records in French. LREC 2026 – 15th biennial Language Resources and Evaluation Conference, May 2026, Palma, Spain. pp.57-65, ⟨10.63317/4cy8xxinjw7z⟩. ⟨hal-05716360⟩
Aygalic Jara-Mikolajczak, Thomas Lavergne, Christophe Servan, Sophie Rosset. Robustesse des LLM dans les contextes longs, hallucinations et détection sur questions-réponses séquentielles. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.512-523. ⟨hal-05708368⟩
Zhongjie Li, Rim Abrougui, Guillaume Lechien, Elisabeth Savatier, Benoît Laurent, et al.. Sem-G-RAG, combiner sémantique symbolique à base de graphes et LLM pour le RAG. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.544-563. ⟨hal-05708371⟩
Eve Sauvage, Cyril Grouin, Julien Tourille. Tous les tokens sont-ils utiles pour les modèles de langues ?. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.595-609. ⟨hal-05708373⟩
Oumaima El Khettari, Virgile Barthet, Guillaume Hocquet, Joconde Weller, Emmanuel Morin, et al.. Is Clinical Text Enough? A Multimodal Study on Mortality Prediction in Heart Failure Patients. LREC 2026 – 15th biennial Language Resources and Evaluation Conference, May 2026, Palma, Spain. pp.194-206, ⟨10.63317/47hsfchk79n6⟩. ⟨hal-05709146⟩
Lounès Kebdi, Lubin Longuépée, Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi. Impact de l’affinage de modèles génératifs pour l’inférence en langue naturelle appliquée aux essais cliniques : comparaison avec des approches de *few-shot learning. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.303-321. ⟨hal-05708357⟩
Vincent Claveau, Nicolas Diniz, Juliane Flament, Nihel Kooli, Jose Moreno, et al.. Actes de l’atelier sur l’évaluation des modèles génératifs (LLM) et challenges (EvalLLM 2026). 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. ATALA, 2026. ⟨hal-05708467⟩
Khanh-an C. Quan, Camille Guinaudeau, Shin’Ichi Satoh. Évaluation de la cohérence des modèles vision-langage pour la tâche de question-réponse visuelle. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.64-74. ⟨hal-05708471⟩
You Zuo, Kim Gerdes, Éric de la Clergerie, Benoît Sagot. Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval. CORIA-TALN 2026 – 21e Conférence en Recherche d’Information et Applications (CORIA), Jun 2026, Nantes, France. ⟨hal-05707237⟩
Younes Djemmal, Olutola Oloruntobi Paul, Kim Gerdes. Au-delà des résumés : Apprentissage des représentations d’articles scientifiques à partir de fenêtres de texte intégral. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.31-50. ⟨hal-05708484⟩
Marc Evrard, Rémi Uro, Nicolas Hervé, Béatrice Mazoyer. French Tweet Corpus for Automatic Stance Detection. Proceedings of the Twelfth International Conference on Language Resources and Evaluation (LREC 2020), ELRA, May 2025, Marseille, France. pp.6317-6322, ⟨10.63317/3kp6n9sci4au⟩. ⟨hal-05682655⟩
Lorena de la Garza, Julie Halbout, Julie Lascar, Niels Martinez-Guevara, Arturo Curiel, et al.. Extracting Signs from Weakly Aligned Sign Language Corpora: A Study on LSF and LSM. 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion (LREC 2026), May 2026, Palma De MaJorque, Spain. pp.174-183, ⟨10.63317/38kfot52b4dz⟩. ⟨hal-05688151⟩
Lucía Catalán Gris, Kim Gerdes. On the difficulty of producing good linguistic lies. ARTS – Atelier sur l’Analyse et la Recherche de Textes Scientifiques 2026, Jun 2026, Nantes, France. ⟨hal-05688347⟩
Idrissa Mahamoudou Dicko, Nona Naderi. Synergizing Domain-Specific Masked Language Models and Instruction-Tuned LLMs for Chemical NER. Atelier IAIntelligence Artificielle et santé, Jun 2026, Arras, France. ⟨hal-05679910⟩