Nicolas Obin

I am Associate Professor at Sorbonne University and a senior research scientist in the Sound Analysis and Synthesis team at the Sound and Music Sciences and Technologies laboratory (Ircam, CNRS, Sorbonne University, French Ministry of Culture).

My background is primarily in mathematics, computer science, and physics, including a Master internship at the CMNAT, University of Berkeley, Califronia under the supervision of Adrian Freed and David Wessel and graduate of the 2005-2006 Master 2 ATIAM (Acoustics, Signal Processing and Computer Science Applied to Music) at the Université Pierre et Marie Curie; secondarily in musicology, with a Master 2 in Arts, Philosophy and Aesthetics from the Université Vincennes Saint-Denis in 2006 under the supervision of Ivanka Stoïanova. I have a doctoral thesis in computer science and telecommunications entitled: ‘MeLos: analysis and modelling of speech prosody and speaking style’ (2011) under the supervision of Xavier Rodet, for which I was awarded the prize for the best doctoral thesis by the Fondation Des Treilles in 2011. In 2023, I defended my Habilitation à Diriger des Recherches (HDR) entitled: ‘From signal modelling to representation learning: structured modelling of speech signals’.  As part of the PostGenAI@Paris AI cluster (France 2030, 2025–2029), I currently co-lead the AI-MADE research program, which focuses on developing next-generation neural generative models for sound and audio creation. The program aims to advance expressivity, data and computational efficiency, and controllability in generative AI through an artist-in-the-loop approach that integrates creative practitioners throughout the entire research and development process, from model conception to artistic uses.

At the crossroads of life sciences, signal and information processing, machine learning, and creative practices, my research focuses on the modeling of behavior and communication among humans, animals, and artificial systems—particularly through sound. Sound lies at the core of my research. I am interested in how it articulates signal and symbol, matter and meaning, communication and creation. I develop generative modeling approaches for complex dynamical systems arising from expressive human behaviors such as speech, singing, and music. I investigate how these systems can be understood, simulated, and extended, and how they can give rise to new forms of sonic and multimodal generation. This research is situated within the broader framework of human and musical cyber-physical systems, as well as emerging forms of creative practice that are increasingly assisted and augmented by artificial intelligence.

More specifically, my research investigates neural approaches to audio generation, with a particular focus on intuitive, controllable and expressive sound generation. I work on autoregressive and temporally synchronized models for audio generation, as well as on learning structured representations in latent spaces. I investigate interpretable sound representations in order to bridge low-level signal structure with higher-level perceptual and cognitive representations, while enabling generative models to interpolate, adapt, and maintain coherence within latent spaces. These approaches support personalized and interactive forms of generation, driven by users and centered on artists. Finally, my work also addresses efficient generation under resource constraints, including limited computational budgets and low-data regimes where high-quality annotated data is scarce. Overall, these approaches aim to enable fine-grained, meaningful control over generation while preserving alignment with perceptual, behavioral, and expressive dimensions of sound.

Beyond research, I engage deeply with the technological, creative, and ethical implications of these systems. In particular, I question how artificial intelligence is reshaping artistic creation, creative professions, and the cultural and creative industries. My research further addresses issues of inclusion and digital sovereignty, with a focus on preserving the diversity of dialects, languages, and cultures in digital environments. As part of my artistic practice at IRCAM, I work at the frontier between research and creation, contributing to the development and dissemination of digital science and technology for the arts and culture. This commitment is grounded in ongoing collaborations with internationally renowned artists across music, cinema, and sound design.

Current functions: 


News


Supervision de thèses de doctorat

Encadrement de thèse (en cours)

[ 2025-2028 ]  Anthony Gallien, Machine Learning for Acoustical In-Painting in Augmented Reality: Enhancing Immersive Audio Realism, École doctorale informatique, télécommunications et électronique (EDITE). Bourse doctorale du Sorbonne Cluster for Articial Intelligence (SCAI). Direction et co-encadrement avec Markus Noisternig et Benoit Alary (STMS, équipe EAC).

[ 2024-2027 ] Balthazar Bujard, Modèles de couplage entre signaux temporels pour le contrôle créatif de la synthèse sonore, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Frédéric Bevilacqua (Direction) et Jérôme Nika (STMS, équipe ISMM).

[ 2024-2027 ] Diego Andres Torres Guarrin, Conversion neuronale des attributs de la voix, projet ANR EVA, École doctorale informatique, télécommunications et électronique (EDITE). Direction et co-encadrement avec Axel Roebel (STMS).

[2023-2026] Téo Guichoux, Génération multimodale du comportement et transfert de style pour l’animation
d’un agent virtuel, bourse du ministère, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Laure Soulier (Direction) et Catherine Pelachaud (ISIR)

[ 2023-2026 ] Mathilde Abrassart, Lightweight Neural Models for Voice Conversion using Parametric Speech Representations, projet ANR BRUEL, École doctorale informatique, télécommunications et électronique (EDITE). Co-direction avec Axel Roebel (Direction, STMS).

Encadrement de thèse (soutenue)

[ 2023-2026 ] Théodor Lemerle, Toward Expressive and Long-form Speech Synthesis, projet ANR EXOVOICES, École doctorale informatique, télécommunications et électronique (EDITE). Co-direction avec Axel Roebel (Direction, STMS). Thèse soutenue le 29 juin 2026.

[2017-2025] Lisa La Pietra, Fonction et approches de la vocalité lyrique et contemporaine aujourd’hui. L’interprétation vocale entre le Belcanto et les nouvelles technologies. Co-encadrement avec Antonio Lai (Direction, Université Vincennes -- Saint-Denis). École doctorale «Esthétique, sciences et technologies des arts» (EDESTA). Thèse soutenue le 9 décembre 2025.

[ 2019-2022 ] Clément Le Moine, Neural conversion of social attitudes in speech signals, en collaboration avec Stellantis, programme doctoral Ph2D/IDF, École doctorale informatique, télécommunications et électronique (EDITE).  Co-encadrement avec Axel Roebel (Direction, STMS). Thèse soutenue le 27 février 2023

[ 2019-2022 ] Mireille Fares, Multimodal expressive gesturing with style, programme doctoral AI @ Sorbonne Université, École doctorale informatique, télécommunications et électronique (EDITE). Co-encadrement avec Catherine Pelachaud (Direction, ISIR). Thèse soutenue le 15 février 2023

[ 2019-2022 ] Killian Martin, Cognitive control of Rooks’ vocalizations, ED 549 Santé, Sciences Biologiques et Chimie du Vivant, Université de Tours, 2019. Co-encadrement avec Valérie Dufour (Direction, CNRS). Thèse soutenue le 13 décembre 2022.

[ 2013-2016 ] Olivier Migliore, Analyser la prosodie musicale du punk, du rap et du ragga français (1977-1992)
à l’aide de l’outil informatique, co-encadrement avec Yvan Nommick (Direction), École doctorale Langues, littératures, cultures, civilisations, Université Montpellier 3. Participation à l'encadrement. Thèse soutenue le 13 décembre 2016.


Le projet TheVoice sélectionné pour les 20 ans de l'ANR

𝟮𝟬 𝗮𝗻𝘀 | 𝟮𝟬 𝘀𝗰𝗶𝗲𝗻𝘁𝗶𝗳𝗶𝗾𝘂𝗲𝘀 | 𝟮𝟬 𝗽𝗿𝗼𝗷𝗲𝘁𝘀 | 𝟮𝟬 𝗿𝗲𝗴𝗮𝗿𝗱𝘀 𝘀𝘂𝗿 𝗹𝗮 𝗿𝗲𝗰𝗵𝗲𝗿𝗰𝗵𝗲 Depuis 2005, l’Agence nationale de la recherche soutient la recherche dans toute sa diversité. En 20 ans, plus de 32 000 projets ont ainsi été financés par l’ANR. Et autant d’histoires et d’aventures scientifiques et humaines. La série de portraits #monANR revient sur ce que ces projets ont changé dans la vie des scientifiques, et sur l’impact de leurs recherches sur la société.


GRABUGE -- Groupe de Recherche sonore et Autres Bidouilles Utopiques, Géniales et Éphémères (2025)

G.R.A.B.U.G.E est un espace de rencontres, d’échanges, et d’expérimentation ouvert à tous les étudiants passionnés de son, de musique, de danse, de réalités mixtes, de machines et autres geekeries phoniques et sensibles ! Les principes de notre démarche : le bricolage, l’expérimentation, l’auto-organisation, la convivialité, et l’entraide. En un mot, un joyeux chaos organisé pour se réunir et faire de la musique avec des machines !

GRABUGE


Un projet soutenu par l'Alliance Sorbonne Université dans le cadre du projet « SOUND - pour un nouvel engagement » financé par l'ANR au titre de France 2030 (ANR-22- EXES-0004) et par l’Union européenne NextGenerationEU.


Tribune : Pour une intelligence artificielle responsable au service d’une création musicale inventive et diverse (2024)


L'IA au service du sonore ? UNESCO (2024)

Soirée "L'IA au service du sonore?" 18 janvier 2024 Organisée dans le cadre de la 21ème édition de la semaine du son

Nicolas Obin, conférence de presse, UNESCO


Soutenance d'habilitation à diriger des recherches

Nicolas Obin soutient son Habilitation à Diriger des Recherches (HDR) le 12 septembre 2023 à 14h - "De la représentation du signal à l’apprentissage de représentations : modélisation structurée de signaux de parole »

Composition du jury

• M. Thomas HUEBER, Directeur de recherche CNRS, GIPSA lab, Rapporteur
• M. Emmanuel VINCENT, Directeur de recherche INRIA, MultiSpeech, Rapporteur
• M. Bjorn SCHULLER, Professeur, Imperial College London, Rapporteur
• M. Gérard BIAU, Professeur, Sorbonne Université, Examinateur
• M. Jean-François BONASTRE, Directeur de Recherche INRIA, Défense et Sécurité, Examinateur
• Mme Catherine PELACHAUD, Directrice de recherche CNRS, ISIR, Examinatrice
• M. Axel ROEBEL, Directeur de recherche, IRCAM, Examinateur
• Mme Isabel TRANCOSO, Professeure, INESC - Université de Lisbonne, Examinatrice
• Mr Nicolas BECKER, Designer sonore et artiste, Membre Invité

Le texte de mon HDR est librement accessible sur HAL.

Publications

In link with

En poursuivant votre navigation sur ce site, vous acceptez l'utilisation de cookies pour nous permettre de mesurer l'audience, et pour vous permettre de partager du contenu via les boutons de partage de réseaux sociaux. En savoir plus.