I am an Associate Professor at Sorbonne University and a senior research scientist in the Sound Analysis and Synthesis team at the Sound and Music Sciences and Technologies Laboratory (STMS), a joint research unit of IRCAM, CNRS, Sorbonne University, and the French Ministry of Culture.
My background combines mathematics, computer science, and musicology. I obtained my PhD in Computer Science and Telecommunications in 2011, with a thesis entitled MeLos: Modelling Prosody and Speaking Style for Text-to-Speech Synthesis, supervised by Xavier Rodet. In 2023, I defended my Habilitation à Diriger des Recherches (HDR), entitled From Signal Representation to Representation Learning: Structured Modelling of Speech Signals. During my studies, I also worked at the Center for New Music and Audio Technologies (CNMAT) at the University of California, Berkeley.
My research focuses on the analysis, modelling, and generation of expressive speech and sound, at the intersection of machine learning, signal processing, and artistic creation. Sound lies at the core of my research, which explores how expressive behaviours can be captured, represented, and generated in speech, singing, and music. I work on structured and interpretable representations, controllable and reactive generation, and efficient generative models to enable fine-grained and meaningful control while aligning with human perception and artistic intentions. I am particularly interested in generative modelling in low-data regimes and under limited computational resources. This research has evolved from the analysis and representation of expressivity in signals toward the development of expressive generative models in which expressivity itself becomes a property to be learned, represented, and generated.
As part of the PostGenAI@Paris AI cluster (France 2030, 2025–2029), I currently co-lead the AI-MADE research program at IRCAM, which focuses on next-generation neural generative models for sound and audio creation. The program explores expressivity, controllability, and data and computational efficiency through an artist-in-the-loop approach, integrating creative practitioners into the design, development, and evaluation of generative models.
Beyond the development of generative models, my research addresses the technological, creative, and ethical implications of artificial intelligence for artistic creation, creative professions, and cultural and creative industries. I am particularly interested in how AI can support new forms of augmented creation, enabling artists to explore, shape, and interact with generative systems while preserving their creative agency and intentions. I also work on questions of inclusion and digital sovereignty, including the preservation of linguistic, dialectal, and cultural diversity in increasingly AI-mediated digital environments. As part of my artistic practice at IRCAM, I work at the interface between research and creation, contributing to the development and dissemination of digital science and technology for the arts and culture. This commitment is grounded in ongoing collaborations with internationally renowned artists across music, cinema, and sound design.
Friday July, 24th: Tell me all. Tell me now sur un texte de James Joyce (Philippe Manoury, world premiere), for voice and electronics with Juliana Snapper (mezzo), Academia Chigiana in Sienne (Italy). Super-resolution audio restoration of historical James Joyce's recordings.
Chairofthe Artificial Intelligence and Databases (IA-BD)Committee, Doctoral School of Computer Science, Telecommunications and Electronics (EDITE),2026DoctoralFellowshipAwardCampaign
Monday, June 1, 2026 – “Scientific Evidence and the Voice”, organized by the Association Francophone de la Communication Parlée, Palais de la Bourse – Lyon, France
Supervision de thèses de doctorat
Encadrement de thèse (en cours)
[ 2025-2028 ] Anthony Gallien,Machine Learning for Acoustical In-Painting in Augmented Reality: Enhancing Immersive Audio Realism, École doctorale informatique, télécommunications et électronique (EDITE). Bourse doctorale du Sorbonne Cluster for Artificial Intelligence (SCAI). Direction et co-encadrement avec Markus Noisternig et Benoit Alary (STMS, équipe EAC).
[ 2024-2027 ] Balthazar Bujard, Modèles de couplage entre signaux temporels pour le contrôle créatif de la synthèse sonore, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Frédéric Bevilacqua (Direction) et Jérôme Nika (STMS, équipe ISMM).
[ 2024-2027 ] Diego Andres Torres Guarrin, Conversion neuronale des attributs de la voix, projet ANR EVA, École doctorale informatique, télécommunications et électronique (EDITE). Direction et co-encadrement avec Axel Roebel (STMS).
[2023-2026] Téo Guichoux, Génération multimodale du comportement et transfert de style pour l’animation d’un agent virtuel, bourse du ministère, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Laure Soulier (Direction) et Catherine Pelachaud (ISIR)
[ 2023-2026 ] Mathilde Abrassart, Lightweight Neural Models for Voice Conversion using Parametric Speech Representations, projet ANR BRUEL, École doctorale informatique, télécommunications et électronique (EDITE). Co-direction avec Axel Roebel (Direction, STMS).
Encadrement de thèse (soutenue)
[ 2023-2026 ] Théodor Lemerle, Toward Expressive and Long-form Speech Synthesis, projet ANR EXOVOICES, École doctorale informatique, télécommunications et électronique (EDITE). Co-direction avec Axel Roebel (Direction, STMS). Thèse soutenue le 29 juin 2026.
[2017-2025] Lisa La Pietra, Fonction et approches de la vocalité lyrique et contemporaine aujourd’hui. L’interprétation vocale entre le Belcanto et les nouvelles technologies. Co-encadrement avec Antonio Lai (Direction, Université Vincennes -- Saint-Denis). École doctorale «Esthétique, sciences et technologies des arts» (EDESTA). Thèse soutenue le 9 décembre 2025.
[ 2019-2022 ] Clément Le Moine, Neural conversion of social attitudes in speech signals, en collaboration avec Stellantis, programme doctoral Ph2D/IDF, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Axel Roebel (Direction, STMS). Thèse soutenue le 27 février 2023
[ 2019-2022 ] Mireille Fares, Multimodal expressive gesturing with style, programme doctoral AI @ Sorbonne Université, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Catherine Pelachaud (Direction, ISIR). Thèse soutenue le 15 février 2023
[ 2019-2022 ] Killian Martin, Cognitive control of Rooks’ vocalizations, ED 549 Santé, Sciences Biologiques et Chimie du Vivant, Université de Tours, 2019. Co-encadrement avec Valérie Dufour (Direction, CNRS). Thèse soutenue le 13 décembre 2022.
[ 2013-2016 ] Olivier Migliore, Analyser la prosodie musicale du punk, du rap et du ragga français (1977-1992) à l’aide de l’outil informatique, co-encadrement avec Yvan Nommick (Direction), École doctorale Langues, littératures, cultures, civilisations, Université Montpellier 3. Participation à l'encadrement. Thèse soutenue le 13 décembre 2016.
Le projet TheVoice sélectionné pour les 20 ans de l'ANR
𝟮𝟬 𝗮𝗻𝘀 | 𝟮𝟬 𝘀𝗰𝗶𝗲𝗻𝘁𝗶𝗳𝗶𝗾𝘂𝗲𝘀 | 𝟮𝟬 𝗽𝗿𝗼𝗷𝗲𝘁𝘀 | 𝟮𝟬 𝗿𝗲𝗴𝗮𝗿𝗱𝘀 𝘀𝘂𝗿 𝗹𝗮 𝗿𝗲𝗰𝗵𝗲𝗿𝗰𝗵𝗲
Depuis 2005, l’Agence nationale de la recherche soutient la recherche dans toute sa diversité. En 20 ans, plus de 32 000 projets ont ainsi été financés par l’ANR. Et autant d’histoires et d’aventures scientifiques et humaines. La série de portraits #monANR revient sur ce que ces projets ont changé dans la vie des scientifiques, et sur l’impact de leurs recherches sur la société.
GRABUGE -- Groupe de Recherche sonore et Autres Bidouilles Utopiques, Géniales et Éphémères (2025)
G.R.A.B.U.G.E est un espace de rencontres, d’échanges, et d’expérimentation ouvert à tous les étudiants passionnés de son, de musique, de danse, de réalités mixtes, de machines et autres geekeries phoniques et sensibles !
Les principes de notre démarche : le bricolage, l’expérimentation, l’auto-organisation, la convivialité, et l’entraide.
En un mot, un joyeux chaos organisé pour se réunir et faire de la musique avec des machines !
Un projet soutenu par l'Alliance Sorbonne Université dans le cadre du projet « SOUND - pour un nouvel engagement » financé par l'ANR au titre de France 2030 (ANR-22- EXES-0004) et par l’Union européenne NextGenerationEU.
Tribune : Pour une intelligence artificielle responsable au service d’une création musicale inventive et diverse (2024)
L'IA au service du sonore ? UNESCO (2024)
Soirée "L'IA au service du sonore?"
18 janvier 2024
Organisée dans le cadre de la 21ème édition de la semaine du son
Soutenance d'habilitation à diriger des recherches
Nicolas Obin soutient son Habilitation à Diriger des Recherches (HDR) le 12 septembre 2023 à 14h - "De la représentation du signal à l’apprentissage de représentations : modélisation structurée de signaux de parole »
Composition du jury
• M. Thomas HUEBER, Directeur de recherche CNRS, GIPSA lab, Rapporteur • M. Emmanuel VINCENT, Directeur de recherche INRIA, MultiSpeech, Rapporteur • M. Bjorn SCHULLER, Professeur, Imperial College London, Rapporteur • M. Gérard BIAU, Professeur, Sorbonne Université, Examinateur • M. Jean-François BONASTRE, Directeur de Recherche INRIA, Défense et Sécurité, Examinateur • Mme Catherine PELACHAUD, Directrice de recherche CNRS, ISIR, Examinatrice • M. Axel ROEBEL, Directeur de recherche, IRCAM, Examinateur • Mme Isabel TRANCOSO, Professeure, INESC - Université de Lisbonne, Examinatrice • Mr Nicolas BECKER, Designer sonore et artiste, Membre Invité
Le texte de mon HDR est librement accessible sur HAL.
En poursuivant votre navigation sur ce site, vous acceptez l'utilisation de cookies pour nous permettre de mesurer l'audience, et pour vous permettre de partager du contenu via les boutons de partage de réseaux sociaux. En savoir plus.