some flowers

Théo Charlot

PhD student @ ComLearn, Grenoble — speech, vision & language acquisition

prof_pic.jpg

Somewhere between Grenoble and its mountains

theocyrilcharlot@gmail.com

I am a PhD student in Grenoble, in the ComLearn team (joint between GIPSA-lab and Inria Grenoble), supervised by Thomas Hueber, Stéphane Lathuilière and Laurent Girin. I work on how multimodal models can learn language the way children do, by looking around and listening to noisy everyday life rather than reading text. I’m broadly interested in human-inspired machine learning, developmentally-plausible learning, and human reasoning.

Before this, I was a Research Engineer at École Normale Supérieure (ENS Ulm), working with the CoML and LAAC teams on speech representation learning and child-directed speech detection in naturalistic, long-form recordings. I completed a Master’s degree in Computer Science (Apprentissage et Traitement Automatique de la Langue) at Nantes Université, graduating with Highest Honors, after a Bachelor’s in Computer Science at Sorbonne Université. Along the way I’ve interned at Inria Paris (ALMAnaCH team) and ENS Ulm (CoML team).

news

Oct 01, 2026 🎓 Starting my PhD in Grenoble, in the ComLearn team, joint between GIPSA-lab and Inria Grenoble, supervised by Thomas Hueber, Stéphane Lathuilière and Laurent Girin, on how multimodal models can learn language the way children do, from perception (audio and video) and social interaction rather than text. How do children learn language by looking around and listening to their noisy everyday life ?
Jun 04, 2026 🎉 Two papers accepted at Interspeech 2026 on tools for child-centered long-form analysis: BabyHuBERT, a multilingual speech representation model for segmenting speakers (Who speaks when?), and addressee, on context-aware child-directed speech detection (Who speaks to whom?).
Nov 03, 2025 🔬 Started as a Research Engineer at École Normale Supérieure (ENS Ulm), CoML & LAAC teams, fine-tuning BabyHuBERT for child-directed speech classification of adult speech. Who speaks to whom ?
Apr 07, 2025 🔬 Started my M2 research internship at École Normale Supérieure (ENS Ulm), CoML team, pretraining and fine-tuning BabyHuBERT, a speech representation model for segmenting speakers in child-centered long-form recordings. Who speaks when ?
May 27, 2024 🌱 Started my M1 research internship at INRIA Paris, ALMAnaCH team, evaluating LLMs’ ability to retrieve common ground information in long dialogs.

selected publications

  1. Context-aware child-directed speech detection from long-form recordings
    Théo Charlot*, Tarek Kunze*, Kaveri K. Sheth, Alejandrina Cristia, and Marvin Lavechin
    In Interspeech, 2026
  2. Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
    Biswesh Mohapatra, Théo Charlot, Giovanni Duca, Mayank Palan, Laurent Romary, and Justine Cassell
    In Findings of the Association for Computational Linguistics, 2026
  3. BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
    Théo Charlot*, Tarek Kunze*, Maxime Poli, Alejandrina Cristia, Emmanuel Dupoux, and Marvin Lavechin
    In Interspeech, 2026