BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//wp-events-plugin.com//7.4.1//EN
TZID:Europe/Paris
X-WR-TIMEZONE:Europe/Paris
BEGIN:VEVENT
UID:0-93@lisn.upsaclay.fr
DTSTART;TZID=Europe/Paris:20230209T110000
DTEND;TZID=Europe/Paris:20230209T120000
DTSTAMP:20230208T162043Z
URL:https://www.lisn.upsaclay.fr/evenements/on-the-use-of-semantically-ali
 gned-speech-representations-for-spoken-language-understanding-with-cross-l
 ingual-and-cross-modal-specialization/
SUMMARY:On the Use of Semantically-Aligned Speech Representations for Spok
 en Language Understanding\, with Cross-lingual and Cross-modal Specializ
 ation
DESCRIPTION:Spoken language understanding (SLU) refers to natural language 
 processing tasks related to semantic extraction from speech.&nbsp\;Differe
 nt tasks can be addressed as SLU tasks\, such as named entity recognition 
 from speech\, call routing\, slot filling task in a context of human-machi
 ne dialogue.&nbsp\;\nEnd-to-end neural approaches have been proposed in or
 der to directly extract the semantics from speech signal\, by using a sing
 le neural model\, instead of applying a classical cascade approach based o
 n the use of an automatic speech recognition (ASR) system\, followed by a 
 natural language understanding processing (NLU) module.&nbsp\;A main issue
  of these approaches is the lack of bimodal annotated data (speech audio r
 ecordings with semantic manual annotation). &nbsp\;Self-supervised learnin
 g (SSL)\, that benefits from unlabelled data\, has been successfully appli
 ed to several SLU tasks\, especially through cascade approaches.\nThe use 
 of an end-to-end approach exploiting directly both speech and text SSL mod
 els is limited by the difficulty to unify the speech and textual represent
 ation spaces\, in addition to the complexity of managing a huge number of 
 model parameters.\nEarlier in 2022\, a new promising model was introduced:
 &nbsp\;SAMU-XLSR (https://arxiv.org/abs/2205.08180)\, for Semantically-Ali
 gned Multimodal Utterance-level Cross-Lingual Speech Representation learni
 ng framework.&nbsp\;The model combines a state-of-the-art multilingual aco
 ustic frame-level speech representation learning model XLS-R with the Lang
 uage Agnostic BERT Sentence Embedding (LaBSE) model to create an utterance
 -level multimodal multilingual speech encoder.\nWe analyzed the performanc
 e and the behavior of the SAMU-XLSR model using the French MEDIA benchmark
  dataset\, which is considered as a very challenging benchmark for SLU.&nb
 sp\;Moreover\, by using the Italian PortMEDIA corpus\, we also investigate
  the potential of porting an existing end-to-end SLU model from one langua
 ge (French) to another (Italian) through two scenarios: zero-shot and low-
 resource learning.\nWe then investigated the use of cross-lingual and cros
 s-modal representations applied to our SLU task. We searched for a way to 
 exploit the sentence-level embeddings produced by SAMU-XLSR in order to im
 prove the semantics extraction and we proposed an efficient specialization
  of the SAMU-XLSR model to sentences related to our task.&nbsp\;In additio
 n\, we presented the benefits of the proposed approach in a language porta
 bility purpose.
CATEGORIES:STL
LOCATION:LISN Site Belvédère\, Rue du Belvédère  Campus universitaire b
 ât 507 91405 Orsay\, France
X-APPLE-STRUCTURED-LOCATION;VALUE=URI;X-ADDRESS=Rue du Belvédère  Campus 
 universitaire bât 507 91405 Orsay\, France;X-APPLE-RADIUS=100;X-TITLE=LIS
 N Site Belvédère:geo:0,0
END:VEVENT
BEGIN:VTIMEZONE
TZID:Europe/Paris
X-LIC-LOCATION:Europe/Paris
BEGIN:STANDARD
DTSTART:20221030T020000
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
TZNAME:CET
END:STANDARD
END:VTIMEZONE
END:VCALENDAR