Mathieu Fontaine is an Associate Professor in Machine Listening at Télécom Paris , affiliated with the LTCI Lab within the IDS Department (Information, Data, Signal). His research focuses on machine listening for speech and audio signal processing. PhD in Informatics (2019), Lorraine University Master in Applied and Fundamental Mathematics (2015), Poitiers University BSc in Fundamental Mathematics (2013), Rennes University Fontaine's research spans speech enhancement , speaker separation , source localization , and music source separation using heavy-tailed probabilistic models and deep Bayesian networks , with applications in augmented reality . He has expertise in Python , signal processing , and machine learning (80% proficiency). His recent publications (2024) include work on diffusion models for speech synthesis , room acoustics estimation from 3D meshes , robust audio scene analysis , and direction-aware speech processing . Earlier publications (2022-2023) explore flow-based NMF , alpha-stable representations , and adaptive beamforming in multiparty environments. Fontaine collaborates with the S2A team and ADASP group at LTCI Lab. His work integrates probabilistic modeling with deep learning to address challenges in real-world audio processing, including reverberation, noise, and complex acoustic environments.
Hao-Wen Dong is an Assistant Professor in the Department of Performing Arts Technology at the University of Michigan, with an affiliation to the Computer Science and Engineering Department. His research focuses on Human-Centered Generative AI for content creation, emphasizing music, audio, and video domains. He holds a Ph.D. in Computer Science from UCSD, advised by Julian McAuley and Taylor Berg-Kirkpatrick. Affiliations: University of Michigan (Primary), UCSD (Ph.D.), National Taiwan University (B.S.) Research Pillars: Generative AI models for new domains, AI-assisted creative tools, and multimodal content creation His work spans music generation (e.g., MuseGAN), audio synthesis (e.g., ViolinDiff), and multimodal systems (e.g., TeaserGen). He has led over 25+ publications in top venues like ISMIR, ICASSP, and ICLR. He advises students in interdisciplinary projects and teaches courses on AI Music and Generative AI for Music/Audio Creation. Notable awards include the Doctoral Award for Excellence in Research (2024) and Rising Stars in AI (2024).
Laxmidhar Behera is a Professor in the Department of Electrical Engineering at the Indian Institute of Technology Kanpur, specializing in Intelligent Systems and Control. With over two decades of academic experience at IIT Kanpur and international research experience at institutions including Fraunhofer Institute of Autonomous Intelligent Systems in Germany, ETH Zurich, and University of Ulster, he has established himself as a leading researcher in cognitive robotics and intelligent control systems. Dr. Behera's research spans multiple cutting-edge domains including Cognitive Robotics, Nano-robotics, Vision based Control, Soft Computing, Information Retrieval in music and language, Semantic Information Processing, Physics of Complex Systems, Cyber Physical Systems, Formation Control of UAVs, Brain-Computer Interface (BCI), and Sanskrit Computational Linguistics. His interdisciplinary approach bridges traditional control theory with modern computational intelligence techniques, creating innovative solutions for complex real-world problems. His extensive publication record in top-tier journals like IEEE Transactions demonstrates his leadership in areas such as brain-computer interfaces, visual servoing, multi-robot systems, and music information retrieval. Notably, his work on quantum neural networks for EEG filtering and multisatellite formation control has received significant attention in the research community. UKIERI Standard Research Award 2008 Best Paper at International Conf. on Intelligent Sensors and Information Processing (ICISIP-2004) Best Paper at WoSco,02, Int. Conf. High-Performance Computing (HiPC, 2002) AICTE career award for young teacher (1997) Senior Member IEEE Multiple IEEE top accessed articles (2009-2010) As an Associate Editor for Autosoft Journal and Technical Committee Member for Intelligent Control at IEEE Control System Society, Dr. Behera actively contributes to the academic community. His laboratory in the Western Lab - 212A of the Department of Electrical Engineering serves as a hub for research in intelligent systems, where he mentors students and collaborates with researchers worldwide on cutting-edge projects in robotics, control systems, and computational intelligence.
Shoji Makino is a Professor at Waseda University's Graduate School of Information, Production and Systems. He has held academic and research positions at institutions such as the University of Tsukuba and NTT Communication Science Laboratories. His work spans acoustic signal processing, blind source separation, and adaptive filtering. Education: Ph.D., Tohoku University (1993.03) Mechanical Engineering, Tohoku University Graduate School of Engineering (1979.04–1981.03) Engineering, Tohoku University Faculty of Engineering (1975.04–1979.03) Research Interests: His research focuses on acoustic signal processing for speech and audio, including blind source separation (BSS) , beamforming , and adaptive filtering . He pioneered methods for solving permutation alignment in frequency-domain BSS and developed geometrically constrained ICA techniques. Scientific Awards: Hoko Award (2018.10, Hattori Hokokai Foundation) Outstanding Contribution Award of the Institute of Electronics, Information, and Communication Engineers (2018.06) IEEE Signal Processing Society Best Paper Award (2014.01) IEEE Fellow (2004.01) IEICE Achievement Award (1997.05) Committee Memberships: He has served as Chair of the IEEE CAS Society's Blind Signal Processing TC, General Chair of IEEE WASPAA2007, and Associate Editor of IEEE Trans. SAP. He is actively involved in EURASIP, APSIPA, and the Acoustical Society of Japan.
Alessandro Ragano is a Postdoctoral Researcher at the Insight Centre for Data Analytics , where he has been investigating Quality of Experience (QoE) aspects of audio archives and developing data-driven approaches for QoE estimation and audio restoration using deep learning since 2018. Education: MSc in Computer Science and Engineering from Politecnico di Milano (Italy) BSc in Computer Engineering from Università Degli Studi di Salerno (Italy) His research integrates machine learning , audio signal processing , and multimedia quality assessment to improve speech enhancement, audio restoration, and perceptual modeling. Recent trends in his publications focus on self-supervised learning , objective quality metrics , and audio dataset generation with applications in speech separation, music representation, and audio inpainting. He actively contributes to open-source tools like Binamix and AQP for audio research and quality evaluation.
Cheng Zhi Huang is the Robert N. Noyce Career Development Professor and Assistant Professor at MIT, holding a shared appointment between the departments of Music and Theater Arts and Electrical Engineering and Computer Science (EECS). His work bridges artificial intelligence, music technology, and computer science to advance human-AI collaboration in musical creativity. Huang leads research in generative models for music composition, real-time interactive systems, and expressive performance synthesis. His contributions include tools like ReaLJam for AI-assisted jamming and the MAESTRO dataset for piano performance modeling. His research interests span AI-driven music generation, human-AI interaction frameworks, and culturally-aware music technologies. Notable projects include The Bach Doodle—an accessible web-based composition tool—and MIDI-DDSP for detailed performance control. Huang’s work emphasizes ethical and creative applications of AI in arts, fostering collaborations between musicians, engineers, and computer scientists. His publications highlight advancements in hierarchical generative modeling, source separation techniques, and co-creation interfaces for novices. Huang’s research has been showcased in venues like TISMIR and IEEE conferences, reflecting his interdisciplinary impact on music technology and machine learning.
Magdalena Fuentes is an Assistant Professor of Music Technology and Integrated Design & Media at New York University (NYU), affiliated with the Music and Audio Research Lab (MARL) and the Integrated Design & Media (IDM) programs. She holds a Ph.D. from Université Paris Saclay (France) and a B.Eng. in Electrical Engineering from Universidad de la República (Uruguay). Her research focuses on Machine Listening, Human-Centered Machine Learning, and Multimodal Representation Learning, with applications to Music Information Retrieval and Environmental Sound Analysis. Previously, she served as a Postdoctoral Faculty Fellow at NYU’s MARL and Center for Urban Science and Progress (CUSP). Her work bridges technical innovation with cultural and societal contexts, particularly in underrepresented music traditions like Brazilian percussion and Candomblé rituals. She has developed open-source tools such as Soundata to ensure reproducible audio dataset usage. Her research outputs span audio-visual synchronization (e.g., SONIQUE), environmental sound monitoring (e.g., SONYC projects), and rhythm analysis frameworks (e.g., Carat toolbox). These contributions highlight her dual focus on advancing machine learning techniques while maintaining human-centric and culturally sensitive applications.
Josh Reiss is a Professor of Audio Engineering at Queen Mary University of London (QMUL), part of the School of Electronic Engineering and Computer Science . He holds additional roles including President-Elect and Fellow of the Audio Engineering Society (AES), and Visiting Professor at Birmingham City University. His research focuses on audio signal processing, procedural audio, and intelligent music production. He earned degrees including BSc in Physics, BSc in Mathematics, and a PhD. Research & Awards : Reiss has published over 200 papers, authored books like Intelligent Music Production , and received awards such as the AES Board of Governors Award (2009, 2010) and Best JAES Paper 2016. His work spans sound synthesis, dynamic range compression, and live audio systems. Teaching & Industry : Teaches modules like Artificial Intelligence and Sound Design. Co-founded startups LandR (AI mixing), Tonz, and Nemisindo. Leads the Centre for Digital Music at QMUL, advancing research in audio technology. Labs & Teams : Active in the Centre for Digital Music, collaborating on projects like the Open Multitrack Testbed and semantic audio evaluation tools.
Dr. Johan Pauwels is a Lecturer in Audio Signal Processing at Queen Mary University of London's School of Electronic Engineering and Computer Science, where he is affiliated with the Centre for Digital Music and the Centre for Multimodal AI. His educational background includes: Master of Science in Electrical/Electronics Engineering from KU Leuven (2006) Master of Science in Artificial Intelligence from KU Leuven (2007) PhD from Ghent University (2016) on automatic harmony recognition from audio Johan's research focuses on making machines understand audio to the level of a trained professional. His work combines machine learning, signal processing, data science, and music theory to develop tools for musicians, listeners, and music learners. He has been working on narrowing the gap between academic research and user-centric applications, web-based music services, and the personalization of spatial and immersive audio. His specific interests include machine learning for audio, audio signal processing, music information retrieval, and binaural audio. His recent publications show a strong focus on music representation learning, with particular attention to limited data scenarios, multimodal approaches, and spatial audio processing. His work bridges theoretical music concepts with practical machine learning applications, especially in chord recognition, beat detection, and instrument recognition. He has made significant contributions to HRTF (Head-Related Transfer Function) research and development of tools for spatial audio processing. Dr. Pauwels is actively involved in research funding, with current grants including the AIM CDT Internship with Sofilab (2025), AIM CDT Studentship - Stem (2024), and AIM CDT Internship with stem.tech (2024). He currently supervises multiple PhD students, primarily through the UKRI Doctoral School in AI and Music, with research topics spanning intelligent audio editing, neural drum synthesis, source separation, graph neural networks for music recommendation, and more. In addition to PhD supervision, he typically guides 8-10 undergraduate and 8-10 master's students through their final year projects. His teaching responsibilities include ECS7013P Deep Learning for Audio and Music (MSc/PhD level) and ECS411U Signals and Information (first-year undergraduate).
Gaël Richard is a Professor at Télécom Paris specializing in machine learning and audio signal processing. He leads the Hi! Paris center, focusing on AI and data science applications. His research emphasizes hybrid interpretable AI for sound analysis, including projects like Hi-Audio funded by a €2.5M ERC Advanced Grant (2022). Key areas include machine listening, music source separation, and speech processing. Applications span autonomous vehicle acoustics and music technology. Notable contributions include neural audio compression (QINCODEC), diffusion models for music synthesis (Diff-TONE), and source separation techniques (Inverse Drum Machine). Research & Awards Recipient of the 2022 ERC Advanced Grant for the Hi-Audio project exploring hybrid AI models that integrate domain knowledge with neural networks. This approach reduces data requirements and enhances model interpretability. Active in audio-visual scene analysis and weakly-supervised learning systems. Affiliations & Labs Executive Director of Hi! Paris, a multidisciplinary lab advancing AI and data science for societal impact. Collaborates on projects like the HI-AUDIO online platform for distributed music data collection and the MAD-EEG EEG dataset for auditory attention decoding.
Ryan Corey (he/him/his) is an Assistant Professor at the University of Illinois Chicago in the Department of Electrical and Computer Engineering, College of Engineering. His research focuses on audio and acoustic signal processing for human listening technologies such as hearing aids and augmented reality systems. Education: MS in Electrical and Computer Engineering from University of Illinois at Urbana-Champaign (2014), BSE in Electrical Engineering from Princeton University (2012) Teaching: Co-instructor for TE 401: Develop Breakthrough Projects, ranked as excellent by students multiple times Corey's research explores signal processing algorithms that combine audio signals from distributed microphone arrays to isolate desired sounds from background noise. His work has potential applications in hearing devices that can enhance specific speaker voices in crowded environments. He also investigates innovative prototyping approaches using unconventional microphone placements on wearable devices. His recent publications focus on advanced beamforming techniques, adaptive filtering methods, and sensor network applications in audio processing. These include topics such as random projections in beamforming, neural adaptive filters, and sound field interpolation using acoustic sensor networks. Scientific Honors: Best Student Paper Award at WASPAA 2019 Intelligence Community Postdoctoral Fellowship (2019) Microsoft Research Dissertation Grant (2018) National Science Foundation Graduate Research Fellowship (2013) Corey leads the Listening Technology Lab and collaborates with Professor Andrew Singer in the Augmented Listening Laboratory. He has mentored students in developing innovative listening technologies, including prototypes for hearing aids with improved directional audio capabilities and devices that modify sound environments for better hearing experiences.
Anna Huang is an Assistant Professor at the Massachusetts Institute of Technology (MIT), affiliated with the PI Core/Dual program. Her research focuses on AI-driven music technologies, including human-AI collaboration, generative music models, and interactive creative tools. She specializes in developing frameworks for real-time music jamming, adaptive accompaniment systems, and novice-friendly AI co-creation platforms. Huang has contributed to projects like the Bach Doodle and the AI Song Contest , demonstrating scalable applications of machine learning in music composition. Her work bridges computer science and musicology, with a particular emphasis on cross-cultural music generation (e.g., Hindustani classical music modeling) and expressive control mechanisms for generative systems. Key areas include MIDI signal processing, source separation algorithms, and the design of user interfaces that empower both professionals and novices to co-create with AI. Huang’s publications emphasize interdisciplinary innovation, with trends spanning reinforcement learning for music performance, hierarchical generative modeling, and ethical considerations in AI-assisted creativity. Though no awards are explicitly listed, her impactful projects suggest recognition in computational music research. Her research also involves dataset development (e.g., MAESTRO dataset) and open-source tools like Coconet, fostering reproducibility and community engagement in music technology.
Richard M. Dansereau is a Professor in the Department of Systems and Computer Engineering at Carleton University's Faculty of Engineering and Design. He holds a Ph.D. from the University of Manitoba and is a Professional Engineer (P.Eng.) and Senior Member of IEEE. He currently serves as Associate Dean (Graduate Studies) and Clerk of Senate, reflecting his leadership roles within the university. His research interests include: Multimodal and audio-visual signal processing Biomedical and biometric signal processing Image and speech signal processing Compressive sensing and deep learning for reconstruction Fractal and multifractal complexity measures, including Rényi dimensions Applications in medical imaging, speech enhancement, and radar systems Recent publications highlight his lab's focus on advanced deep learning techniques for image reconstruction (e.g., deep equilibrium models for compressive sensing), medical image analysis (e.g., PET reconstruction and cervical cell segmentation), and Riemannian geometry in radar signal processing for drone detection. His work integrates theoretical signal processing with practical applications in healthcare and defense. Scientific awards associated with his research group include: Ontario Graduate Scholarship Alexander Graham Bell Canada Graduate Scholarship (CGS D) John Ruptash Memorial Fellowship NSERC Best Project Award 1st prize in poster competition at hSITE 2012 Finalist for World Congress Award at WSCTS’2006 Dansereau actively supervises graduate students, with a long list of Ph.D. and M.A.Sc. alumni who have worked on topics such as speech separation, ECG analysis, image registration, and radar signal processing. He collaborates with researchers at institutions like the University of Ottawa Heart Institute and Defence Research and Development Canada (DRDC). His lab, the Signal Processing and Machine Learning Lab, continues to publish in top journals and conferences, securing research opportunities for Canadian, American, and British citizens in speech intelligibility research.
George Tzanetakis is a Professor in the Department of Computer Science at the University of Victoria, Canada, affiliated with the Faculty of Engineering and Computer Science. He holds a PhD from Princeton University. His research focuses on computer audition, audio signal processing, machine learning, and music information retrieval, with applications in human-computer interaction, music education, and robotics. He is involved in the Music and Sound Interdisciplinary Centre, exploring innovative technologies for music performance, analysis, and interaction. Key areas of exploration include real-time gesture-based control systems, robotic musical instruments, and audio-based health monitoring. His work integrates multimodal data (e.g., audio, motion) and leverages deep learning to address challenges in music technology, bioacoustics, and healthcare analytics. Recent projects include developing frameworks for audio representation evaluation, modular synthesis systems, and PU learning models for healthcare prediction. His research has led to advancements in music transcription, sound event detection, and interactive musical interfaces. Collaborations span academia and industry, emphasizing practical applications of computational methods in music and beyond. No specific awards or grants are listed in the provided text.
Pauline Larrouy-Maestri is a Senior Researcher at the Max Planck Institute for Empirical Aesthetics in Frankfurt/Main, Germany, where she has been working since 2019 after serving as a Postdoctoral Researcher in the Neuroscience Department from 2014-2019. Her interdisciplinary research focuses on how humans categorize acoustic information that unfolds over time to make sense of sounds, working at the intersection of music, speech, and neuroscience. Dr. Larrouy-Maestri holds a PhD in Psychology from the University of Liège (2009-2013) and has an unusually diverse educational background including a Bachelor in Music (Piano) from the Royal Conservatory of Mons, a Master in Speech Therapy from the University of Brussels, additional studies in Psychology, Pedagogy, and Music Therapy, and research stays at McGill University and SUNY Buffalo. This multidisciplinary foundation informs her unique approach to studying sound perception. Her research examines how we process ambiguous auditory material that sits at the boundaries between music and speech categories, such as sprechgesang and West-African talking drums. She investigates auditory sequence processing in music, particularly how continuous streams of sound are parsed into meaningful units, and has made significant contributions to understanding the perception of correctness in singing. Her work on vocal communication explores how pitch, timing, and other acoustic features contribute to our interpretation of emotional content and meaning in both music and speech. Analysis of her recent publications reveals a sophisticated integration of behavioral, electrophysiological, and computational approaches to study music-speech interactions, with growing emphasis on cross-cultural perspectives, individual differences, and neural mechanisms. Her work increasingly examines how subtle acoustic variations influence aesthetic judgments and emotional responses to vocalizations. 2023: €20,000 research scholarship for "Humanity of Speech" project 2017: Selected for "Sign Up! Careerbuilding for outstanding female post docs in the MPG" 2016: Young Investigator Award from SEMPRE and ICMPC14 2015: PBEEE Merit scholarship from Fonds de recherche du Québec 2013: Patrimoine de l'Université de Liège and FNRS fundings 2011: Grant from French Community of Belgium Dr. Larrouy-Maestri currently supervises multiple researchers including Camila Bruder, Madita Hoerster, and Zofia Hobubowska. Her research is supported by competitive grants including the recent Imminent scholarship and previous funding from Belgian and Canadian sources. She maintains extensive international collaborations with researchers including David Poeppel, Melanie Wald-Fuhrmann, Marc Pell, and others across neuroscience, psychology, and musicology disciplines. Her work is conducted within the Neuroscience Department at the Max Planck Institute for Empirical Aesthetics, where she contributes to the institute's interdisciplinary mission of studying aesthetic experiences through multiple methodological approaches. She participates in research groups focusing on auditory perception, music cognition, and the neural mechanisms underlying language and music processing, helping bridge traditionally separate fields through innovative experimental designs.