Jan Østergaard is a Full Professor in Information Theory and Signal Processing at Aalborg University's Department of Electronic Systems. He leads the AI and Sound research section and directs the CASPR center. His expertise spans AI-driven acoustic signal processing, information theory, and EEG signal analysis. Østergaard holds a M.Sc. from Aalborg University and a PhD (cum laude) from Delft University of Technology. Major awards include the Danish Young Researcher’s Award and a EURASIP Best Thesis honor. His work focuses on speech enhancement, sound zone technologies, and neural tracking of auditory attention. Recent research emphasizes low-latency speech transmission, deep learning for sound field control, and robust voice activity detection. He serves on editorial boards and national committees, advancing Denmark’s sound technology initiatives. Education: M.Sc. (Aalborg, 1999), PhD (Delft, 2007) Research interests emphasize practical AI applications in sound systems, including hearing aid improvements, data-efficient acoustic modeling, and feedback control in networked systems. Over 210 publications and 17 active projects reflect his interdisciplinary impact across academia and industry.
Sara Ng is a Visiting Assistant Professor in the Department of Linguistics at Western Washington University. She holds a PhD in Linguistics from the University of Washington (2024) and an MS in Computational Linguistics (2023), alongside a B.A. in Linguistics and B.S. in Applied Mathematics from the University of Utah (2017). Her research focuses on computational models of prosody, speech perception, and their integration with speech technologies like automatic speech recognition (ASR). She investigates how prosody conveys pragmatic meaning, influences conversational dynamics, and impacts listeners with hearing impairments. Her work bridges computational methods and linguistic theory, addressing challenges in clinical and technological applications. Education PhD in Linguistics, University of Washington, 2024 MS in Computational Linguistics, University of Washington, 2023 B.A. in Linguistics (Honors), B.S. in Applied Mathematics, University of Utah, 2017 Research Interests Ng’s research explores computational linguistics, prosody modeling, speech technology, and hearing impairment studies. She develops methods to leverage prosodic cues for tasks like punctuation prediction and investigates the impact of hearing loss on speech perception. Grants & Awards 2024: Nominated for UW Excellence in Teaching Award 2023: Excellence in Linguistic Research Graduate Fellowship 2017: University of Utah Top Scholar Award Teaching Ng teaches courses in phonetics, computational linguistics, and linguistics for honors students. She emphasizes pedagogical innovation and inclusivity, with experience as an instructor of record and teaching assistant at both Western Washington University and the University of Washington. Labs & Collaborations She collaborates with the TIAL Lab (University of Washington), Phonetics Lab, Hearing Aid Laboratory (Northwestern University), and CLILLAC-ARP (Université Paris Cité). Her work intersects with clinical and engineering domains, addressing real-world applications of speech technology.
Lerrel Pinto is an Assistant Professor of Computer Science at the Courant Institute of Mathematical Sciences at New York University (NYU), where he leads the General-purpose Robotics and AI Lab (GRAIL) as part of the CILVR research group. His work bridges the gap between theoretical machine learning and practical robotics applications, with a focus on enabling robots to generalize and adapt in real-world environments. Dr. Pinto received his undergraduate degree from IIT Guwahati, followed by a PhD from the Robotics Institute at Carnegie Mellon University (CMU). He then completed a postdoctoral fellowship at the University of California, Berkeley before joining NYU as faculty. His research program centers on robot learning and decision making, with several key thrusts that demonstrate his innovative approach to robotics. Pinto's work emphasizes large-scale learning techniques that leverage both extensive data and sophisticated model architectures. A significant portion of his research focuses on representation learning for sensory data, particularly developing methods that enable robots to make sense of visual, tactile, and auditory inputs. His lab has made notable contributions to reinforcement learning algorithms that allow robots to adapt to new scenarios with minimal retraining. Pinto also champions open-source robotics , developing affordable robot platforms that democratize access to robotics research. Analysis of Pinto's recent publications reveals a strong trend toward multimodal perception in robotics, integrating visual, tactile, and auditory information to create more robust robot systems. His work increasingly focuses on zero-shot and few-shot learning capabilities, enabling robots to handle novel situations without extensive retraining. There's also a clear progression toward general-purpose robotics , moving away from task-specific solutions toward more flexible systems that can handle diverse real-world challenges. Dr. Pinto's scientific contributions have been recognized with several prestigious awards: Sloan Research Fellowship (2025) NSF CAREER Award (2024) RAL Early Career Award (2024) Best Student Paper Award at ICRA (2016) Outstanding Paper Award at MFM-EAI workshop at ICML (2024) Best Paper Award at NGSM workshop at ICML (2024) Best Student Paper Award at RSS (2023) As an advisor, Pinto has mentored numerous students who have gone on to impactful careers in both academia and industry. His former PhD student Denis Yarats co-founded Perplexity.AI, while Mahi Shafiullah became a postdoc at UC Berkeley and Meta AI. Many of his Masters students have pursued PhDs at top institutions like CMU, MIT, and Stanford, or joined leading robotics companies including 1X, Fauna Robotics, and NVIDIA. Pinto's lab has secured significant research funding, including the NSF CAREER award and likely other grants supporting his robotics research program. The General-purpose Robotics and AI Lab (GRAIL) that Pinto leads brings together a diverse team of researchers working on cutting-edge robotics challenges. The lab maintains strong collaborations with industry partners and other academic institutions, facilitating technology transfer and real-world impact. GRAIL's research spans multiple robotics platforms and focuses on developing algorithms that enable robots to learn from diverse experiences and generalize across environments.
Brent Doiron is a Professor at the University of Chicago, holding appointments in the Departments of Neurobiology and Statistics, and serving on the Committee on Computational and Applied Mathematics (CCAM). His research integrates nonlinear dynamics and statistical mechanics to study neural circuit variability, focusing on mechanisms underlying neural coding and network learning through collaborations with experimentalists in sensory systems. Education: PhD in Physics (University of Ottawa, 2004) Postdoc: Center for Neural Science at New York University (2017) Previous Roles: Mathematics Professor at University of Pittsburgh (2007-2020), Co-Director of Neural Computation Program at Carnegie Mellon Neuroscience Institute Research interests center on neuronal population dynamics, recurrent circuit mechanisms, and computational neuroscience. Current work investigates correlated variability in cortical networks, inter-areal communication, and stochastic spiking models. Recent publications emphasize cortical stability/gain modulation, asynchronous/synchronous activity balance, and Bayesian inference frameworks. Key themes include sensory processing, network plasticity, and dimensionality reduction in neural coding. Scientific Awards Alfred P. Sloan Research Fellowship in Neuroscience Vannevar Bush Faculty Fellowship Chancellor’s Distinguished Research Award (University of Pittsburgh) Active grants include NIH R01 and R90/T90 awards for neuronal dynamics research and computational neuroscience training programs.
Professor Bernd Möbius is a leading academic in Phonetics and Phonology at the Department of Language Science and Technology, Saarland University. His research bridges phonetic theory with speech technology applications, focusing on text-to-speech systems, prosody modeling, and computational simulations of speech processes. Current research projects: DFG SFB 1102, C1: Information density and phonetic structure predictability DFG SFB 1102, C4: Slavic intercomprehension and surprisal theory (INCOMSLAV) Research Themes: Key areas include text-to-speech synthesis, speech prosody analysis, experimental methods in speech production/perception, information density in phonetics, and cross-linguistic studies of Slavic-Germanic languages. Scientific Contributions: Recent work explores Parkinson-induced dysarthria detection, breath noise acoustics, surprisal-driven speech behaviors, multilingual BERT models for idiomaticity, and perceptual consequences of acoustic adjustments.
Olga Sorkine-Hornung is a Professor of Computer Science at ETH Zurich and head of the Institute of Visual Computing. She leads the Interactive Geometry Lab, focusing on theoretical and practical advancements in digital content creation, geometry processing, and shape modeling. Current position: ETH Zurich, Department of Computer Science Previous roles: Courant Institute (NYU), Technical University of Berlin Education: BSc and PhD from Tel Aviv University, postdoc at TU Berlin Her research spans shape representation, digital fabrication, computer animation, and fundamental geometry processing. Key contributions include Laplacian surface editing, as-rigid-as-possible deformation, and generalized winding numbers. She works on applications in VR, AR, and autonomous systems. Awards include Test of Time Awards (2024), ACM Fellow (2020), ERC Consolidator Grant (2020), and EUROGRAPHICS Young Researcher Award (2008). She has supervised numerous students and co-developed software libraries like libigl and Instant Meshes . Co-chair roles for SIGGRAPH, Eurographics, and Pacific Graphics Editorial board member for ACM Transactions on Graphics and other journals Keynote speaker at VMV, CVPR, and SIAM conferences Her work bridges mathematical rigor with practical implementation, advancing computer graphics and geometry processing through intuitive algorithms that maintain surface detail while enabling efficient computation.
Nicolas Mathevon is a Professor at the University of Saint-Etienne and a Senior Member of the Institut universitaire de France (IUF). His research focuses on bioacoustics and animal communication, spanning from marine mammals to birds. He has held visiting positions at Hunter College (City University of New York) and the University of California, Berkeley. Education Habilitation à diriger les Recherches (2002, University of Saint-Etienne) PhD (1996, University of Lyon 1) Master's Degree (ranked 2nd, University of Lyon 1) Agrégation in Life and Earth Sciences (ranked 8th) Mathevon's research explores how animals use sound for social interactions, mate selection, and environmental adaptation. His work bridges bioacoustics, neuroethology, and behavioral ecology, with a focus on decoding vocal signals' complexity. His publications cover diverse topics including animal soundscapes, vocal memory in seals, and neural encoding of communication calls. The articles highlight interdisciplinary approaches combining neuroscience, behavioral studies, and acoustic signal analysis. Scientific Awards Prose Award for Excellence in Biological and Life Sciences, AAP (2024) Chevalier des Palmes Académiques (2017) IUF Senior Member (2015-present) IUF Junior Member (2005-2010) Mathevon has supervised 17 PhD students, 10 postdocs, and contributed to public outreach through media, documentaries, and public lectures. He leads the International Master of Bioacoustics program and serves as President of the International BioAcoustic Society.
Andrés Buxó-Lugo serves as an Assistant Professor of Psychology at the University at Buffalo, where he directs the Language Processing and Computation Lab. His research investigates the cognitive mechanisms underlying language production, comprehension, and acquisition with a specialized focus on speech prosody—the rhythm, intonation, and intensity patterns in speech—and their role in human communication. His primary research interests include psycholinguistics, cognitive psychology, speech prosody, language production, language comprehension, language acquisition, and computational linguistics. He examines how listeners integrate diverse linguistic cues during speech processing, how individuals learn unfamiliar constructions like non-native pronunciations or novel prosodic patterns, and the cognitive basis of durational changes in speech. His work also explores how communicative context shapes prosodic production and how higher-level linguistic information aids prosodic structure parsing. Analysis of his 15 most recent publications (2019-2025) reveals consistent interdisciplinary work bridging cognitive science, linguistics, and computational modeling. Key trends include phonological representation studies, speech planning mechanisms, intonation adaptation across talkers, lexical representation structures, and the integration of input expectations in syntactic parsing. His research demonstrates significant methodological diversity, incorporating experimental paradigms, computational modeling, and acoustic analysis to unravel language processing complexities. As director of the Language Processing and Computation Lab at the University at Buffalo, Buxó-Lugo leads research initiatives focused on developing computational models of language processing while investigating the cognitive foundations of speech and prosody through empirical experimentation and theoretical innovation.
Pasquale Bottalico serves as Associate Professor in the Department of Speech and Hearing Science at the University of Illinois, with dual appointments as Associate Professor at the Center for Latin American and Caribbean Studies and Affiliate Faculty in the School of Music. His unique interdisciplinary profile bridges engineering, music performance, and speech science, reflecting his dual academic training and professional artistry. His educational foundation includes: Bachelor's in Telecommunications Engineering from Univeristà Mediterranea di Reggio Calabria, Italy Concurrent Opera Singing degree from F. Cilea Music Academy, Reggio Calabria Master's in Telecommunications Engineering from Politecnico di Torino, Italy Ph.D. in Metrology specializing in acoustics measurement uncertainty and classroom acoustics Dr. Bottalico's research centers on vocal load quantification and professional voice techniques , with significant contributions to understanding vocal fatigue in teachers and singers. His work spans Speech Intelligibility in educational environments, Room Acoustics for performance and learning spaces, and Musical Acoustics of historical vocal styles. A distinctive thread throughout his research examines how acoustic conditions modulate voice production and perception, increasingly incorporating virtual reality and bone conduction technologies for innovative assessment and intervention approaches. His Colombian vocal health study demonstrates cross-cultural applications of his work. Analysis of his 2023-2025 publications reveals three dominant research trajectories: (1) The impact of noise and dysphonia on children's speech processing in educational settings, using multimodal assessment including EEG; (2) Virtual reality applications for voice production research and therapeutic intervention; (3) Cross-cultural validation of vocal fatigue metrics and development of biofeedback systems. His work consistently bridges engineering precision with clinical applicability, particularly for professional voice users in challenging acoustic environments. No scientific awards were documented in the available information. While specific advising relationships aren't detailed, his research collaborations span international institutions including Colombian and Italian universities, suggesting graduate mentorship in interdisciplinary projects. No grant information was provided, though his systematic reviews and cross-cultural studies imply externally funded research activities. Though no dedicated laboratory is specified, his virtual reality voice studies and acoustic parameter assessments suggest affiliations with audio engineering facilities and voice clinics, likely through the Speech and Hearing Science department's research infrastructure.
Xavier Serra is a Full Professor at the Department of Engineering at Universitat Pompeu Fabra (UPF), Barcelona. He is the founder and director of the Music Technology Group (MTG), and leads the UPF-BMAT Chair on AI and Music. He also coordinates the Master in Sound and Music Computing and serves as President of the Phonos Foundation. His research focuses on audio signal processing, sound and music computing, and computational musicology, emphasizing open science and open innovation. Education: BSc in Biology, University of Barcelona (1981) Master in Music, Florida State University (1983) PhD in Computer Music, Stanford University (1989) Research Interests: Audio Signal Processing Data-Driven and Knowledge-Driven Methodologies Music Information Retrieval Cultural Music Analysis (e.g., Carnatic/Turkish/Andalusian Music) Music Education Technology Notable Projects: CompMusic (ERC Advanced Grant, 2010-2017): Multicultural computational music analysis Open datasets: Freesound, Saraga, FSD50K Technologies: Reactable, Vocaloid, Essentia API Recent Trends in Articles: Focus on AI-driven audio processing (neural fingerprints, generative models), cross-cultural music analysis, and explainable music difficulty estimation. Awards: ERC Advanced Grant (2010) for CompMusic Project. Labs/Teams: Director of MTG, Phonos Foundation, and UPF-BMAT Chair. Active in open-source projects and international collaborations.
Nicolas Mathevon is a Professor at the University of Saint-Etienne, with a focus on Animal Behavior and Bioacoustics . He currently serves as Director of Studies at Ecole Pratique des Hautes Etudes and leads the ENES Team . His research explores acoustic communication across diverse species, including crocodiles, seals, penguins, and humans, with particular emphasis on vocal recognition, signal evolution, and neurobiological underpinnings. Major collaborations include Dr. T. Aubin (CNRS), Dr. I. Charrier (marine mammals), Dr. N. Grimault (crocodilian bioacoustics), and Prof. D. Reby (human communication). He also co-edits a new book, The Voices of Nature , published by Princeton University Press. His recent work spans ecoacoustic monitoring for conservation, multi-modal communication in crocodiles, and cross-species analysis of distress signals in bonobos, chimps, and humans. Neurobiological studies include brain activation patterns in response to vocalizations and sound localization mechanisms in reptiles. Scientific Awards include senior membership in the Institut universitaire de France .
Allison Koenecke is an Assistant Professor of Information Science at Cornell Tech and a field faculty member in Computer Science at Cornell University. Previously, she was a postdoctoral researcher at Microsoft Research New England and completed her PhD at Stanford University’s Institute for Computational & Mathematical Engineering. Her research focuses on algorithmic fairness, computational social science, and causal inference in public health, addressing disparities in automated systems like speech recognition and policy decision-making. Education : PhD, Stanford Institute for Computational & Mathematical Engineering MA/MS, Stanford University BA/BS, Massachusetts Institute of Technology Research Interests : Dr. Koenecke’s work bridges economics and computer science, emphasizing equity in AI systems. Key areas include: Algorithmic bias in speech recognition (e.g., racial disparities in voice assistants) Fairness in policy tools like environmental justice data systems Causal analysis in public health interventions Ethical implications of large language models in education Media & Impact : Her research has been featured in outlets like New York Times , Science , and Scientific American . Notable studies include exposing racial gaps in speech-to-text systems and advocating for inclusive dataset development. She also explores societal impacts of AI in education and governance. Awards : Sloan Research Fellow in Computer Science Forbes 30 Under 30 in Science NSF Awards Cornell CIS Teaching Excellence Award (2024) Teaching & Outreach : She teaches courses like Designing Fair Algorithms and Data Science for Global Development , emphasizing interdisciplinary collaboration. Her PhD Professionalization course addresses hidden curricula in academia. She advises on AI ethics for nonprofits, tech companies, and government agencies.
Jianjing Kuang is an Associate Professor of Linguistics at the University of Pennsylvania, affiliated with the School of Arts and Sciences. As Director of the Penn Phonetics Laboratory, they lead research in phonetics, laboratory phonology, and tonal language studies. Kuang holds a Ph.D. from UCLA (2013) and specializes in the interplay between production and perception in speech, particularly focusing on tonal systems, prosody, and cross-linguistic fieldwork. Their work integrates behavioral experiments, corpus studies, and computational modeling to explore phonological contrasts, voice quality, and sound change. Key affiliations include MindCORE and the Center for East Asian Studies. Research interests include multidimensional cues in tone processing, glottal articulations, mapping production-perception relationships, and prosodic sentence processing. Kuang has conducted fieldwork on languages such as Yi, Q’anjob’al, and Mandarin. Their studies address topics like tonal splitting in Yi, cue-changing in Korean stops, and prosodic patterns in Mayan languages. Recent work explores voice quality’s role in pitch perception and prosodic boundary detection. Awards and grants are not explicitly listed, but Kuang’s extensive publications (over 50 peer-reviewed papers) reflect sustained academic impact. They advise graduate students in phonetics and phonology, with active collaborations in computational linguistics and speech technology. The Penn Phonetics Laboratory, under Kuang’s direction, emphasizes interdisciplinary research bridging experimental and theoretical phonetics.
Nima Mesgarani is an Associate Professor of Electrical Engineering at Columbia Engineering, Columbia University, affiliated with the Sense, Collect and Move Data Committee. His research bridges engineering and neuroscience through reverse-engineering neural signal processing mechanisms, leading to advancements in brain-machine interfaces, neural prosthetics, and speech processing algorithms. He received his PhD in Electrical Engineering from the University of Maryland and completed postdoctoral training at Johns Hopkins University's Center for Language and Speech Processing and UC San Francisco's Neurosurgery Department. Research Focus Professor Mesgarani's lab integrates computational neuroscience and engineering to study acoustic signal processing. Key areas include: Neural decoding of speech and auditory attention in multi-talker environments Development of brain-controlled hearing technologies Novel speech separation and synthesis algorithms inspired by cortical processing Cross-modal learning between auditory and visual systems Applications of large language models in neural signal interpretation Publication Trends Analysis of his 15 most recent articles (2025) reveals dominant themes: neural decoding techniques using intracranial EEG, brain-inspired speech separation models (e.g., Mamba architectures), applications of large language models in auditory neuroscience, cross-modal distillation methods, and clinical translation of audio processing algorithms. A strong emphasis emerges on real-time brain-computer interfaces and noise-robust speech processing. Laboratory and Collaborations Mesgarani directs an interdisciplinary lab developing neurotechnology for hearing restoration. His team collaborates with neurosurgery departments and speech processing centers, focusing on translating theoretical models into clinical brain-machine interfaces. The lab's work has yielded patents for brain-informed speech separation systems and attention-decoding frameworks.
Dhruv Jain is an Assistant Professor in the Computer Science and Engineering Department at the University of Michigan , with affiliations in the School of Information and Michigan Medicine . He leads the Soundability Lab , focusing on transforming hearing into a programmable interface through human-centered AI systems. PhD from University of Washington MS from MIT Media Lab Formerly worked at Microsoft Research, Google, and Apple Research Interests intersect Human-Computer Interaction (HCI) , Audio AI , Accessibility , and Hearing Health . His work develops systems like SoundWatch (smartwatch-based sound awareness) and HomeSound (IoT sound visualization) to expand auditory access for DHH individuals and other domains. Recent Publications (2023–2025) span top venues like CHI , ASSETS , and ICMI , with trends including generative AI for real-time captioning , adaptive soundscapes , and clinical communication tools . SIGCHI Outstanding Dissertation Award (2023) William Chan Memorial Dissertation Award (2023) Best Poster Award at ASSETS 2024 Google Academic Research Award (2024) NIH Grant ($450k, 2023) William Demant Foundation Grant ($740k, 2025) Teaching includes Accessible Computing (undergraduate) and Advanced Accessibility (graduate), alongside global DIY workshops in five countries. The Soundability Lab collaborates with neuroscientists, clinicians, and Deaf scholars to create deployable systems like clinical communication tools now at Michigan Medicine and features integrated into Apple devices.