The Department of Language Sciences and Technologies studies fundamental questions relating to linguistic systems by exploiting large corpora collected, annotated and enriched in an unsupervised or semi-supervised way by statistical learning models adapted to the linguistic material.
These models make it possible to study how languages function, their variations (phonetic-phonological, morphological-lexical, syntactic and semantic), both synchronic and diachronic, diaphasic and diatopic, and to raise questions about their acquisition as mother tongues or second languages. Finally, the department is developing major applications in language processing: speech recognition, automatic translation, information retrieval, conversational agents, etc. … which are increasingly important for society (safeguarding endangered languages, providing tools for people with disabilities, helping to process information and medical knowledge) and for ethics.
This approach to language and languages covers a broad spectrum, from the most fundamental to the most applied research, in a wide variety of media (newspapers, social media, video, telephone, . . .) and all modalities (written, spoken and signed).
This research is highly multidisciplinary, bringing together diverse communities from the fields of computer science, engineering and the humanities.
Ayoub Hammal, Pierre Zweigenbaum, Caio Corro. On the Rejection Criterion for Proxy-based Test-time Alignment. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Jul 2026, San Diego, United States. pp.547-554, ⟨10.18653/v1/2026.acl-short.46⟩. ⟨hal-05689863⟩
Clémentine Bleuze, Karën Fort, Vincent P Martin, Aurélie Névéol. Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024. JMIR AIArtificial Intelligence, 2026, 5, pp.e88082-e88082. ⟨10.2196/88082⟩. ⟨hal-05717650⟩
Marco Naguib, Christel Gérardin, Victor Beaucoté, Cyril Charron, Adrien Joseph, et al.. Evaluating the Retrieval Component in a Retrieval-Augmented Summarization System for Patient Records in French. The Fifteenth Language Resources and Evaluation Conference (LREC 2026), May 2026, Palma, France. pp.57-65, ⟨10.63317/4cy8xxinjw7z⟩. ⟨hal-05716360⟩
Aygalic Jara-Mikolajczak, Thomas Lavergne, Christophe Servan, Sophie Rosset. Robustesse des LLM dans les contextes longs, hallucinations et détection sur questions-réponses séquentielles. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.512-523. ⟨hal-05708368⟩
Zhongjie Li, Rim Abrougui, Guillaume Lechien, Elisabeth Savatier, Benoît Laurent, et al.. Sem-G-RAG, combiner sémantique symbolique à base de graphes et LLM pour le RAG. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.544-563. ⟨hal-05708371⟩
Eve Sauvage, Cyril Grouin, Julien Tourille. Tous les tokens sont-ils utiles pour les modèles de langues ?. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.595-609. ⟨hal-05708373⟩
Oumaima El Khettari, Virgile Barthet, Guillaume Hocquet, Joconde Weller, Emmanuel Morin, et al.. Is Clinical Text Enough? A Multimodal Study on Mortality Prediction in Heart Failure Patients. LREC 2026 – 15th Language Resources and Evaluation Conference, May 2026, Palma, Spain. pp.194-206, ⟨10.63317/47hsfchk79n6⟩. ⟨hal-05709146⟩
Lounès Kebdi, Lubin Longuépée, Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi. Impact de l’affinage de modèles génératifs pour l’inférence en langue naturelle appliquée aux essais cliniques : comparaison avec des approches de *few-shot learning. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.303-321. ⟨hal-05708357⟩
Vincent Claveau, Nicolas Diniz, Juliane Flament, Nihel Kooli, Jose Moreno, et al.. Actes de l’atelier sur l’évaluation des modèles génératifs (LLM) et challenges (EvalLLM 2026). 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. ATALA, 2026. ⟨hal-05708467⟩
Khanh-an C. Quan, Camille Guinaudeau, Shin’Ichi Satoh. Évaluation de la cohérence des modèles vision-langage pour la tâche de question-réponse visuelle. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.64-74. ⟨hal-05708471⟩
You Zuo, Kim Gerdes, Éric de la Clergerie, Benoît Sagot. Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval. CORIA-TALN 2026 – 21e Conférence en Recherche d’Information et Applications (CORIA), Jun 2026, Nantes, France. ⟨hal-05707237⟩
Younes Djemmal, Olutola Oloruntobi Paul, Kim Gerdes. Au-delà des résumés : Apprentissage des représentations d’articles scientifiques à partir de fenêtres de texte intégral. 21e Conférence en Recherche d’Information et Applications (CORIA) 19e Rencontres Jeunes Chercheurs en RI (RJCRI) 33e Conférence sur le Traitement Automatique des Langues Naturelles (TALN) 28e Rencontre des Étudiants Chercheurs en Informatique pour le Traitement Automatique des Langues (RÉCITAL), Jun 2026, Nantes, France. pp.31-50. ⟨hal-05708484⟩
Marc Evrard, Rémi Uro, Nicolas Hervé, Béatrice Mazoyer. French Tweet Corpus for Automatic Stance Detection. Proceedings of the Twelfth International Conference on Language Resources and Evaluation (LREC 2020), ELRA, May 2025, Marseille, France. pp.6317-6322. ⟨hal-05682655⟩
Lorena de la Garza, Julie Halbout, Julie Lascar, Niels Martinez-Guevara, Arturo Curiel, et al.. Extracting Signs from Weakly Aligned Sign Language Corpora: A Study on LSF and LSM. 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion (LREC 2026), May 2026, Palma De MaJorque, Spain. pp.174-183, ⟨10.63317/38kfot52b4dz⟩. ⟨hal-05688151⟩
Lucía Catalán, Kim Gerdes. On the difficulty of producing good linguistic lies. Atelier sur l’Analyse et la Recherche de Textes Scientifiques 2026 (ARTS), Jun 2026, Nantes, France. ⟨hal-05688347⟩
Idrissa Mahamoudou Dicko, Nona Naderi. Synergizing Domain-Specific Masked Language Models and Instruction-Tuned LLMs for Chemical NER. Atelier IAIntelligence Artificielle et santé, Jun 2026, Arras, France. ⟨hal-05679910⟩
Jose Felipe Espinosa Orjuela, Philippe Boula de Mareüil, Marc Evrard. Speech synthesis for Walloon, an under-resourced minority language. 13th edition of the Speech Synthesis Workshop, Aug 2025, Leeuwarden, Netherlands. pp.189-195, ⟨10.21437/SSW.2025-29⟩. ⟨hal-05682647⟩
Benedictus Kent Rachmat, Thomas Gerald, Zheng Zhang, Cyril Grouin. QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings. ACL 2025 – 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), ACL, Jul 2025, Vienna, Austria. pp.1132-1144, ⟨10.18653/v1/2025.acl-srw.89⟩. ⟨hal-05683004⟩
Laura Ascone, Lucie Gianola, Julien Longhi, Laurène Renaut. La linguistique forensique pour l’analyse du discours : anticiper les risques, aider à la décision, répondre aux menaces. Colloque R2DIP, « Les notions de risques, société et sécurité dans les discours institutionnels et politiques », CY Cergy Paris Université, Dec 2017, Cergy, France. ⟨hal-05682141⟩
Théophile Lenoir, Ana Valdivia, Aurélie Bugeau, Anne-Laure Ligozat. Beyond the Energy Efficiency Directive. Observatory on the Environmental Footprint of AIArtificial Intelligence, 2026. ⟨hal-05680820⟩
Iskandar Boucharenc, Eve Sauvage, Thomas Gerald, Julien Tourille, Sabrina Campano, et al.. Using Syntax for the Semantic Representation of Sentences. SLiDE 1st Workshop on Structured Linguistic Data and Evaluation at the 2026 Language Resources and Evaluation Conference (LREC 2026), May 2026, Palma de majorque, Spain. pp.169–179, ⟨10.63317/4gtinxarm3dd⟩. ⟨hal-05669816⟩
Clément Morand, Aurélie Névéol, Anne-Laure Ligozat. The Rising Unsustainability of AIArtificial Intelligence Graphics Cards Production. LIMITS 2026: 12th Workshop on Computing within Limits, Jun 2026, Online, France. ⟨hal-05666542⟩
Kim Gerdes. The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal Dependencies. UDW 2026 – Ninth Workshop on Universal Dependencies, May 2026, Palma De MaJorque, Spain. pp.163-173, ⟨10.63317/4akqrtsv7i65⟩. ⟨hal-05676925⟩
Iskandar Boucharenc, Sahar Ghannay, Christophe Servan, Laure Soulier, Sophie Rosset. Étude de l’adaptation des gros modèles de langues par retour visuel. Journée Visu, GdR IG-RV, Jun 2023, Orsay, France. ⟨hal-05670004⟩
Emmett Strickland, Marc Evrard, Valentina Fedchenko. Transfer Learning for Creole TTS: A Pilot Study on Whether Substrate Phonologies or Lexifier Vocabularies Matter More. Towards Inclusivity and Equality: Language Resources and Technologies for Under-Resourced and Endangered Languages, SIGUL 2026 Joint Workshop with ELE, EURALI, and DCLRL, May 2026, Palma De Majorque, Spain. ⟨10.63317/5d5qjmokuvmc⟩. ⟨hal-05617449⟩
Clémentine Bleuze, Bruno Guillaume, Aurélie Névéol, Karën Fort. Omniprésents et anthropomorphisés : analyse lexico-syntaxique des discours sur les LLM. TALN 2026 – 33e Conférence sur le Traitement Automatique des Langues Naturelles, Jun 2026, Nantes, France. ⟨hal-05670834⟩
Clémentine Bleuze, Karën Fort, Vincent P. Martin, Aurélie Névéol. Grands modèles de langue pour prédire la santé mentale : une revue exploratoire de la documentation des biais et de l’utilité clinique. TALN 2026 – 33e Conférence sur le Traitement Automatique des Langues Naturelles, Jun 2026, Nantes, France. ⟨hal-05670826⟩
Clément Morand, Aurélie Névéol, Rosy Tsopra, Anne-Isabelle Tropeano, Sophie de Chambine, et al.. Prospectively Evaluating the Environmental Impacts of Digital Health Applications : A Case Study and Recommendations. Journal of the American Medical Informatics Association, 2026, ⟨10.1093/jamia/ocag091⟩. ⟨hal-05628404⟩
Thomas Gerald, Sahar Ghannay, Julie Lascar, Paul Lerner, Anne Vilnat. Can Multimodal LLMs Generate Pedagogical Questions?. LREC 2026, May 2026, Palma, Spain. ⟨10.63317/4z4gj3h8jmc7⟩. ⟨hal-05658326⟩
Thierry Hamon. Description of the LISN system for extracting terms. DEfinition and Term Extraction CHallenge 2026 (DETECH 2026), Jun 2026, Zadar, Croatia. ⟨hal-05669893⟩
Marie Schmit, Melvin Selim Atay, Khalid Belhajjame, Ulysse Le Clanche, Emmanuel Coquery, et al.. ShareFAIR-KG, a centralised knowledge base of scientific workflows. JOBIM 2026 – Journées Ouvertes en Biologie, Informatique et Mathématiques, Jun 2026, Strasbourg, France. ⟨hal-05666980⟩
Louis Estève, Marie-Catherine de Marneffe, Nurit Melnik, Agata Savary, Olha Kanishcheva. A survey of diversity quantification in natural language processing: The why, what, where and how. 2026. ⟨hal-05661565⟩
Alexandre Genadot, Nicolas Guilliot, Philippe Boula de Mareüil. Introduction to the book “Cartographier les Langues de Nouvelle-Aquitaine: entre Grammaire et Société”. 2026. ⟨hal-05662837⟩
Agata Savary, Manon Scholivet, Carlos Ramisch, Takuya Nakamura, Eric Bilinski, et al.. PARSEME 2.0 Multilingual Corpus of Multiword Expressions. LREC 2026 – 15th biennial Language Resources and Evaluation Conference, ELRA Language Resources Association, May 2026, Palma De MaJorque, Spain. pp.4819-4834, ⟨10.63317/2iy5qf38yhay⟩. ⟨hal-05661505⟩
Julie Halbout, Annelies Braffort, Michèle Gouiffès, Diandra Fabre, Julie Lascar. Learning to Spot Signs from Named Entities. A study on French Sign Language. LREC 2026 – 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion, May 2026, Palma de Majorque, Spain. ⟨10.63317/26i8n4zuyzyx⟩. ⟨hal-05636077⟩
Damien Lacroux, Aurélie Bugeau, Anne-Laure Ligozat. The indirect rebound effects of AIArtificial Intelligence as undone science: philosophical reflection on two structural causes. Undone Computer Science, Mar 2026, Luxembourg, Luxembourg. ⟨hal-05624399⟩
Benedictus Kent Rachmat, Thomas Gerald, Zheng Zhang, Cyril Grouin. Les données de calibration comptent-elles vraiment pour LoRA?. EvalLLM2026 : Atelier sur l’évaluation des modèles génératifs (LLM), le RAG et challenges, Jul 2026, Nantes (France), France. ⟨hal-05633638⟩
Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi. Assessing the Difficulty of Inference Types in Natural Language Inference for Clinical Trials. The Fifteenth Language Resources and Evaluation Conference (LREC 2026), May 2026, Palma, France. pp.5290-5300, ⟨10.63317/359toazp33g8⟩. ⟨hal-05652719⟩
Jenny Copara, Nona Naderi, Gilles Falquet, Douglas Teodoro. MeSH Concept Relevance and Knowledge Evolution: A Data-Driven Perspective. 12th International Conference on Information Management and Big Data. Communications in Computer and Information Science, Oct 2025, Lima (Pérou), Peru. pp.280-299, ⟨10.1007/978-3-032-20322-9_20⟩. ⟨hal-05625658⟩
Clément Morand, Aina Rasoldier, Paul Gay. Not up to its critical perspective on digitalization: A Descriptive Analysis of How Sustainability is Approached in the ICT4S Conference. ICT4S, Jun 2026, Berne, France. ⟨hal-05615744⟩
Fanny Ducel, Lucie Digoin-Caparros, Ibrahim Al Kotob, Shayan Ahmed Shariff, Binesh Arakkal Remesh, et al.. Les benchmarks sont une source de biais des LLM : MMLU, CommonSenseQA et MGSM au microscope. TALN 2026 – 33e Conférence sur le Traitement Automatique des Langues Naturelles, Jun 2026, Nantes, France. ⟨hal-05618509⟩
Louis Estève, Christophe Servan, Thomas Lavergne, Agata Savary. A Diversity Diet for a Healthier Model: A Case Study of French ModernBERT. 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Jul 2026, San Diego, United States. ⟨hal-05599374⟩
Virgile Barthet. Extraction d’information et classification de textes cliniques pour la prédiction du risque de décès. Intelligence artificielle [cs.AIArtificial Intelligence]. Université Paris-Saclay, 2026. Français. ⟨NNT : 2026UPASG019⟩. ⟨tel-05599487⟩
Luc Pommeret, Thomas Gerald, Christophe Servan, Sahar Ghannay, Patrick Paroubek, et al.. Étude des propositionneurs multilingues : formalisation, évaluation et interprétabilité. CORIA-TALN, ARIA; ATALA, Jun 2026, Nantes, France. ⟨hal-05597666⟩
Manon Scholivet, Agata Savary, Carlos Ramisch, Eric Bilinski, Takuya Nakamura, et al.. Edition 2.0 of the PARSEME shared task on multilingual identification and paraphrasing of multiword expressions. Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), Mar 2026, Rabat, Morocco. pp.254-275, ⟨10.18653/v1/2026.mwe-1.33⟩. ⟨hal-05588684⟩
Jean-Luc Gauvain, Abdel Messaoudi, Holger Schwenk. Language Recognition Using Phone Lattices. International Conference on Speech and Language Processing, Oct 2004, Jeju, South Korea. pp.1283–1286. ⟨hal-01434492⟩
Luc Pommeret, Thomas Gerald, Sophie Rosset, Patrick Paroubek, Christophe Servan, et al.. Les propositions atomiques : un pont entre approches neuronales et symboliques. Journée interprétabilité, GDR TALTraitement Automatique des langues, Mar 2026, Jussieu, Paris, France. ⟨hal-05575718⟩
Luc Pommeret, Thomas Gerald, Patrick Paroubek, Sahar Ghannay, Christophe Servan, et al.. LLM-based Atomic Propositions Help Weak Extractors: Evaluation of a Propositioner for Triplet Extraction. KG-LLM@LREC – Knowledge Graphs and Large Language Models, ELRA, May 2026, Palma De Majorque, Spain. ⟨10.63317/3kna3utavhgb⟩. ⟨hal-05572941⟩
Jean-Luc Gauvain, Gilles Adda, Lori Lamel, Fabrice Lefèvre, Holger Schwenk. Transcription de la parole conversationnelle. Revue TALTraitement Automatique des langues : traitement automatique des langues, 2005, 45 (3). ⟨hal-01434260⟩
Jean-Luc Gauvain, Gilles Adda, Martine Adda-Decker, Alexandre Allauzen, Veronique Gendner, et al.. Where are we in transcribing French broadcast news?. Eurospeech, Sep 2005, Lisbonne, Portugal. pp.1665-1668, ⟨10.21437/Interspeech.2005-544⟩. ⟨hal-01434245⟩
Hélène Bonneau-Maynard, Alexandre Allauzen, Daniel Déchelotte, Holger Schwenk. Combining Morphosyntactic Enriched Representation with n-best Reranking in Statistical Translation. HLT/NACL workshop on Syntax and Structure in Statistical Translation, Apr 2007, Rochester, United States. pp.65-71. ⟨hal-01434104⟩
Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André F T Martins, Ayoub Hammal, et al.. EuroBERT: Scaling Multilingual Encoders for European Languages. COLM 2025 – Second Conference on Language Modeling, Oct 2025, Montreal, Canada. pp.1-28. ⟨hal-05226285⟩
Pierre Lepagnol. Petits modèles génératifs en contexte industriel : Adaptation par prompting avec peu de données. Intelligence artificielle [cs.AIArtificial Intelligence]. Université Paris-Saclay, 2026. Français. ⟨NNT : 2026UPASG011⟩. ⟨tel-05572429⟩
Karin Dassas, Cyrille Bonamy, Bruno Bzeznik, Emmanuelle Frenoux, Gaël Guennebaud, et al.. Estimer l’impact carbone des activités numériques d’une unité de recherche. CNRS (EcoInfo). 2026. ⟨hal-05568070⟩
Ayoub Hammal, Pierre Zweigenbaum, Caio Corro. KAD: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral. EACL 2026 – 19th Conference of the European Chapter of the Association for Computational Linguistics, Mar 2026, Rabat, Morocco. pp.3854-3872, ⟨10.18653/v1/2026.eacl-long.179⟩. ⟨hal-05571208⟩
Clément Morand, Jacques Combaz, Aurélie Névéol, Anne-Laure Ligozat. When rebound effect is not a side effect: analyzing sociotechnical contexts of digital technologies. 2026. ⟨hal-05566029⟩
Natalia Grabar, Cyril Grouin. Year 2021: COVID-19, Information Extraction and BERTization among the Hottest Topics in Medical Natural Language Processing. IMIA Yearbook of Medical Informatics, 2022, 31 (01), pp.254-260. ⟨10.1055/s-0042-1742547⟩. ⟨hal-03931852⟩
Pierre Lepagnol, Sahar Ghannay, Thomas Gerald, Christophe Servan, Sophie Rosset. Format Matters: A Critical Evaluation of Output Formats for Prompting LLMs in SLU and NER. The Fifteenth biennial Language Resources and Evaluation Conference (LREC 2026), May 2026, Palma de Majorque, Spain. ⟨10.63317/3osjjdr778fh⟩. ⟨hal-05546569⟩
Clémentine Bleuze, Fanny Ducel, Maxime Amblard, Karën Fort. COCOA: Creation and Exploratory Investigation of a Corpus of Claims from NLP Articles. LREC 2026 – International Conference on Language Resources and Evaluation, ELRA Language Resources Association, May 2026, Palma de Mallorca, Spain. ⟨10.63317/38hiuxwcq4bc⟩. ⟨hal-05547842⟩
Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi. Assessing the Difficulty of Inference Types in Natural Language Inference for Clinical Trials. 2026. ⟨hal-05533706v2⟩
Juan Manuel Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset, Khaled Zaouk, et al.. Diart: A Python Library for Real-Time Speaker Diarization. Journal of Open Source Software, 2024, 9 (99), pp.5266. ⟨10.21105/joss.05266⟩. ⟨hal-05530961⟩
Clémentine Bleuze, Karën Fort, Vincent P. Martin, Aurélie Névéol. Grands modèles de langue pour la détection de pathologies psychiatriques : promesses, réalité, et enjeux. Journée d’étude “LLM@hopital”, ATALA, Mar 2026, Paris, France. ⟨hal-05532823⟩
Iskandar Boucharenc. Hierarchical Prefixes for Long Document Representations. ECIR – European Conference on Information Retrieval, Apr 2025, Lucca, Italy. pp.171-177, ⟨10.1007/978-3-031-88720-8_28⟩. ⟨hal-05530637⟩
Fanny Ducel, Aurélie Névéol, Vidit Khazanchi, Loïc Leclere, Arthur Pedrini, et al.. Code-switching as a Bias Indicator in LLMs: “The consequences are not the same para nosotros”. LREC 2026 – 15th biennial Language Resources and Evaluation Conference, May 2026, Palma De Mallorca, Spain. ⟨10.63317/2mq6kqjk9bng⟩. ⟨hal-05529786⟩
Oralie Cattan, Christophe Servan, Sophie Rosset. On the Usability of Transformers-based models for a French Question-Answering task. Joint Conference of the Information Retrieval Communities in Europe (CIRCLE) 2022, Jul 2022, Samatan, France. ⟨hal-03701740⟩
Léa Pacini, Jérôme Dupire, Isabelle Barbet, Olivier Pons, Camille Guinaudeau, et al.. Textbook’s accessibility for children with dyspraxia and visual disability. 17th International Conference of the Association for the Advancement of Assistive Technology in Europe, AAATE 2023, Association for the Advancement of Assistive Technology in Europe, Aug 2023, Paris, France. ⟨hal-04410340⟩
Fanny Ducel. How to define, understand and evaluate stereotypical biases in language models?. IASIV – Séminaire du groupe de travail Intelligence Artificielle Sûre, Intelligible et Vérifiable, Mar 2025, Palaiseau, France. ⟨hal-05467784⟩
Gustave Cortal. Natural language processing for subjectivity analysis in personal narratives. Computation and Language [cs.CL]. Université Paris-Saclay, 2026. English. ⟨NNT : 2026UPASG003⟩. ⟨tel-05501345⟩
Julie Halbout, Annelies Braffort, Michèle Gouiffès. Annotation automatique d’un corpus de Langue des Signes Française. RJCP – Rencontres Jeunes Chercheurs en Parole, Nov 2025, Paris, France. ⟨hal-05495878⟩
Annelies Braffort, Michael Filhol, Michèle Gouiffès, Julie Halbout, Julie Lascar. Sign Language Processing with Linguistic Structure. BMVA Symposium on AIArtificial Intelligence for Sign Language Translation, Production, and Linguistics, Dec 2025, London, United Kingdom. ⟨hal-05495664⟩
Jules Françoise, Julie Lascar, Cyril Verrecchia, Sidonie Minodier, Michèle Gouiffès, et al.. LaboSignes : vers une IAIntelligence Artificielle participative pour la reconnaissance automatique de la Langue des Signes Française. Journée d’études AFIA-ATALA : Technologies linguistiques pour les langues peu dotées, Dec 2025, Paris, France. ⟨hal-05495906⟩
S. Rosset, D. Tribout, L. Lamel. Multi-level Information and Automatic dialog Act Detection in Human-Human Spoken Dialogs. Speech Communication, 2008, 50 (1), pp.1-13. ⟨10.1016/j.specom.2007.05.007⟩. ⟨hal-00499189⟩
Idrissa Mahamoudou Dicko, Nona Naderi. Biomedical hallucination detection of LLMs using Med-HALT and HaloScope frameworks. JDSE 2025 – 10th Junior Conference on Data Sciences and Engineering, Sep 2025, Gif-sur-Yvette, France. ⟨hal-05483690⟩
Philippe Boula de Mareüil, Albert Rilliard, Frédéric Vernier. Valorisation de la diversité linguistique à travers un atlas sonore. Myriam Caressa; Christophe Doubovetzky. Langue(s) et droit(s). Enjeux et paradoxes en France, L’Harmattan, pp.177-188, 2025, Logiques Juridiques, 978-2-336-55319-1. ⟨hal-05464189⟩
Natalia Grabar, Thierry Hamon, Emmanuelle Canut. Le langage simplifié pour le public FLE : des critères linguistiques à interroger. Éducation, formation et communication. L’accompagnement des publics en exil. Problèmes de langue et modalités de communication, A paraître, 2865310019. ⟨hal-05465059⟩
Anjani Dhrangadhariya, Roger Hilfiker, Karl Martin Sattelmayer, Nona Naderi, Katia Giacomino, et al.. RoBuster: A Corpus Annotated with Risk of Bias Text Spans in Randomized Controlled Trials in Physiotherapy and Rehabilitation. JMIR Formative Research, 2026, 10, pp.e55127-e55127. ⟨10.2196/55127⟩. ⟨hal-05462769⟩
Fanny Ducel, Karën Fort, Aurélie Névéol. La linguistique appliquée pour une IAIntelligence Artificielle plus éthique. NéALA 2025 – Colloque sur Naturel et Artificiel en Linguistique Appliquée : une époque de paradoxes, Jul 2025, Nancy, France. ⟨hal-05457534⟩
Luciana Benotti, Fanny Ducel, Karën Fort, Guido Ivetta, Zhijing Jin, et al.. Navigating Ethical Challenges in NLP: Hands-on strategies for students and researchers. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 5: Tutorial Abstracts), 2025, ⟨10.18653/v1/2025.acl-tutorials.5⟩. ⟨hal-05457524⟩
Simon Devauchelle, Albert Rilliard, David Doukhan, Lucas Ondel Yang. Variation of Perceived Voice Pitch Across Time Periods, Gender, and Age in French Media Archives. Valentina De Iacovo; Bianca Maria De Paolis; Daniela Mereu. The voice in the media and new technologies, 12 (004), Officinaventuno, pp.47-71, 2024, Studi Associazione Italiana Scienze della Voce, 978-88-97657-73-6. ⟨10.17469/O2112AISV000004⟩. ⟨hal-05450567⟩
Mathieu Laï-King, Patrick Paroubek. Pre-training data selection for biomedical domain adaptation using journal impact metrics. 23rd Workshop on Biomedical Natural Language Processing, Aug 2024, Bangkok, Thailand. pp.363-369, ⟨10.18653/v1/2024.bionlp-1.27⟩. ⟨hal-05447036⟩
Adrien Berthelot, Tiago da Silva Barros, Laurent Lefèvre, Anne-Laure Ligozat, Emeline Pegon. Multi-criteria and multi-stage environmental study of Pl@ntnet service for the year 2024. Inria Lyon. 2026. ⟨hal-05448455v2⟩
François Buet, Camille Guinaudeau, Cyril Grouin, Sahar Ghannay, Shin’ichi Satoh. XAI for Gender Representation in Media Analysis. ICASSP 2025 – 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE Signal Processing Society, Apr 2025, Hyderabad, India. pp.1-5, ⟨10.1109/ICASSP49660.2025.10888945⟩. ⟨hal-05442625⟩
Phrashant Khatri, Hansjörg Mixdorff, Preeti Rao, Albert Rilliard. Recognition of Audio-Visual Attitudes. ESSV – 36. Konferenz Elektronische Sprachsignalverarbeitung, Department of Speech Science and Phonetics of the Institute of Music, Media and Speech Sciences at the Martin Luther University Halle-Wittenberg in Halle/Saale; Central German Association for Speech Science and Speech Education, Mar 2025, Halle / Saale, Germany. pp.19-26. ⟨hal-05426157⟩
Luc Pommeret, Sophie Rosset, Christophe Servan, Sahar Ghannay. AtomicEval: Evaluation Framework for Atomic Proposition Autonomy with French Propositioner. JDSE 2025 – 10th Junior Conference on Data Sciences and Engineering, Sep 2025, Gif-sur-Yvette, France. . ⟨hal-05414939⟩
Michael Filhol. AZVD as a Sign Language writing system proxy, and the potential evolution. Proceedings of Grapholinguistics in the 21st century, Oct 2024, Venice, Italy. ⟨hal-05344585⟩
Bran Knowles, Vicki L Hanson, Christoph Becker, Mike Berners-Lee, Andrew A Chien, et al.. Climate Change: What is Computing’s Responsibility?. 2025, pp.1-18. ⟨10.4230/DagMan.11.1.1⟩. ⟨hal-05369257⟩
Quentin Le Tellier, Marc Evrard, Albert Rilliard, Jean-Sylvain Liénard. Impact de la parole expressive sur l’estimation de l’intensité vocale. CFA 2025 – 17e Congrès Français d’Acoustique, Société Française d’Acoustique (SFA), Apr 2025, Paris, France. ⟨hal-05365670⟩
Jean-Sylvain Liénard, Albert Rilliard, Marc Evrard, Quentin Le Tellier. Variabilité du signal de parole en fonction de la Force de Voix en situation d’interaction orale. CFA 2025 – 17e Congrès Français d’Acoustique, Société Française d’Acoustique (SFA), Apr 2025, Paris, France. ⟨hal-05366097⟩