PHD defence : Théodor Lemerle

  • Research
  • these

Théodor Lemerle, PhD candidate in the EDITE Doctoral School (ED130) at Sorbonne University, conducted his doctoral research entitled “Toward Long-Form and Expressive Speech Synthesis” within the Sound Analysis and Synthesis team of the STMS laboratory (IRCAM, CNRS, Sorbonne University, French Ministry of Culture), under the supervision of Axel Roebel and Nicolas Obin.

This work was carried out as part of the ANR EXOVOICES project, in collaboration with Lunii, the Laboratoire de Sciences Cognitives et Psycholinguistique (LSCP), and IRCAM.

The dissertation defense will be held in English on June 29, 2026, at 2:00 PM in the Stravinsky Room at Ircam. It will be registred on Youtube :  https://youtube.com/live/Z8ZzRfT6MC0

The examination committee will consist of:

Ricard Marxer — Professor, University of Toulon — Reviewer
Geoffroy Peeters — Professor, Télécom Paris, Institut Polytechnique de Paris — Reviewer
Gérard Biau — Professor, Sorbonne University — Examiner
Simon King — Professor, University of Edinburgh — Examiner
Berrak Sisman — Assistant Professor, Johns Hopkins University — Examiner
Alexandre Défossez — Chief Exploration Officer, Kyutai — Examiner
Nicolas Obin — Associate Professor, Sorbonne University — Co-supervisor
Axel Roebel — Research Director, IRCAM — PhD Supervisor

Résumé :

This PhD thesis focuses on neural text-to-speech (TTS) synthesis, and more specifically on its adaptation to expressive story telling. It is conducted within the framework of the ANR EXOVOICES project, in collaboration with Lunii, the Laboratoire de Sciences Cognitives et Psycholinguistique (LSCP), and IRCAM. The advent of neural network-based generative models has led to remarkable progress in speech synthesis, making it possible to generate voices that are difficult to distinguish from human speech. However, these advances rely on increasingly large computational infrastructures and ever-growing training datasets, which predominantly consist of short utterances. As a result, current systems are becoming increasingly costly to reproduce and still struggle to generate long-form narration that is coherent, stable, and expressive. First, we propose a speech synthesis system that enables finer conditioning on stylistic and emotional attributes, along with a method for localized control of these attributes in the absence of specifically annotated training data. Second, we introduce a neural speech codec designed to provide a representation well suited for generation while remaining reproducible on commonly available hardware. Finally, we propose a speech synthesis model capable of continuous generation over arbitrarily long durations without loss of stability or coherence. Our approach stems from an empirical analysis of attention mechanisms in conventional neural speech synthesis systems, which suggests an underutilization of the receptive field. This observation motivates a dedicated windowing strategy, which we show enables synthesis to generalize beyond the training horizon while preserving speaker identity and synthesis quality. Taken together, these contributions advance the development of expressive and controllable speech synthesis systems better suited to long-form narration, while maintaining a reproduction cost compatible with the resources typically available in public research laboratories.

En poursuivant votre navigation sur ce site, vous acceptez l'utilisation de cookies pour nous permettre de mesurer l'audience, et pour vous permettre de partager du contenu via les boutons de partage de réseaux sociaux. En savoir plus.