Andrea Santilli is a Research Scientist at Nous Research and holds a PhD in Computer Science from GLADIA at Sapienza University of Rome. His research focuses on large language models (LLMs), robustness, reliability, and multimodal learning. He previously worked at Apple MLR, Hugging Face’s BigScience, and Pi School. He earned his MSc and BSc in Computer Science from Tor Vergata University and Sapienza. Education: PhD in Computer Science, Sapienza University of Rome (2024) MSc in Computer Science, University of Roma Tor Vergata (2020) BSc in Computer Science, University of Roma Tor Vergata (2018) Research Interests: Santilli’s work spans LLM robustness , mechanistic interpretability , multimodal neural databases , and instruction-tuning . He introduced Parallel Jacobi Decoding and contributed to projects like BLOOM, Camoscio, and Fauno. His research bridges syntax-aware NLP, privacy-preserving LLMs, and cross-modal alignment. Publications: His work includes advancements in 3D-text latent space alignment (CVPR 2025), evolutionary merging (ICML 2025), and efficient decoding (ACL 2023). Over 15+ peer-reviewed papers span venues like ACL, CVPR, and ICLR. Awards: Received the Emanuele Pianta Award for his MSc thesis on continual language learning with syntax-based episodic memory. Grants & Projects: Winner of ‘Machine Learning Algorithms for Translation’ grant (2022), developing Parallel Decoding Co-PI for ‘Multimodal AI for 3D Analysis’ (2021) with Ecole Polytechnique Labs & Teams: Active in GLADIA (Sapienza), Apple MLR, and Hugging Face’s BigScience initiative. Core contributor to open-source projects like PromptSource and BLOOM.
Emily J. King is a tenured Associate Professor in the Department of Mathematics at Colorado State University (CSU), College of Natural Sciences. She previously held a faculty position at the University of Bremen and has been actively contributing to the mathematical community through research, mentorship, and academic leadership. Her primary research interests include Frame Theory , Harmonic Analysis , Algebraic and Geometric Combinatorics , and Data Science , with applications in signal and image processing, Earth science, and artificial intelligence. She integrates deep mathematical theory with practical data analysis challenges. Her recent scholarly output reflects a strong focus on equiangular tight frames, combinatorial structures in frames, mathematical models for attention mechanisms, and applications to satellite imagery and cloud processes. Her work often bridges pure and applied mathematics, with a growing emphasis on interpretable AI and data science foundations. Dr. King has supervised several doctoral and master’s students, including Lander ver Hoef, Sören Schulze, Harley Meade, and Kristina Moen. She is a co-PI on an NSF grant focused on cloud processes and has been recognized for mentoring excellence, as evidenced by her student Emma Slack receiving the inaugural Outstanding Undergraduate in Mathematics award. NSF Grant Co-PI (2024) Outstanding Undergraduate in Mathematics award (mentored student, 2023) She is a founding co-organizer of the international Codes and Expansions (CodEx) Seminar and has organized sessions at major conferences such as SIAM AG and the Joint Mathematics Meetings. She frequently delivers invited talks at universities and research institutes worldwide, including upcoming presentations at the Air Force Institute of Technology, SIAM AG25, and TU Clausthal. Dr. King’s academic lineage includes John Benedetto as her mathematical advisor and Chandler Davis as her mathematical grandfather. She is actively involved in interdisciplinary research, particularly in marine data science, having co-spoken for the Helmholtz School for Marine Data Science (MarDATA).
Tanel Alumäe is an Associate Professor of Speech Processing at Tallinn University of Technology's School of Information Technologies, Department of Software Science. With over 15 years of academic experience, he has held various research and teaching positions at the university since 2006, progressing from Research Fellow to Tenured Associate Professor. His work focuses on speech and language technologies with a particular emphasis on Estonian language applications. PhD in Information and Communication Technology (2006), Tallinn University of Technology Research Master's Degree in Informatics (2002), Tallinn Technical University MSc studies at Tallinn Technical University (1999-2002) and Universität Erlangen-Nürnberg, Germany (1999-2000) Diploma in Computer and Systems Engineering (1994-1999), Tallinn Technical University Alumäe's research spans automatic speech recognition, speaker recognition, natural language processing, and computational linguistics with a focus on Estonian language technology. His work addresses challenges in multilingual speech processing, deep learning applications for speech technologies, and developing practical systems for real-world applications including broadcast media processing and accessibility solutions. He has made significant contributions to low-resource language processing and specialized applications for children's speech and emotion recognition. His recent publications demonstrate a strong focus on cutting-edge speech processing techniques including deepfake detection, multi-speaker systems, speech-to-speech translation, and applying large language models to speech applications. The research shows a consistent pattern of addressing both theoretical challenges in speech processing and practical implementations for Estonian language technology. Award 'Keeletegu 2019' from the Ministry of Education and Research Award 'Keeletegu 2011' from Estonian Ministry of Education and Research 3rd award at the Tallinn University of Technology contest for applied scientific projects (2011) Boris Tamm stipend (2007) First prize at the national contest of students' scientific works (2007) Ustus Agur stipend of Estonian Information Technology and Telecommunications Association (2005) Alumäe has supervised postdoctoral researchers including Rena Nemoto (2012-2015) on pronunciation modeling for speech recognition. He serves in editorial and review capacities for major journals including Nature, Computer Speech & Language, and IEEE Transactions. His administrative roles include Secretary of the Northern European Association for Language Technology Board and membership on the Department of Software Science Council at TalTech. His research group at Tallinn University of Technology actively participates in international challenges (IWSLT, Interspeech, Odyssey) and collaborates with institutions worldwide. The team has developed open-source platforms for Estonian speech transcription and created systems for automatic closed captioning of Estonian broadcasts, demonstrating strong practical applications of their research.
Dr. Swati Chandna is a Senior Lecturer at the School of Computing and Mathematical Sciences, Birkbeck, University of London. She holds an honorary position as an Honorary Lecturer in Statistics at University College London (UCL) from January 2023 to January 2026. She earned her PhD in Statistics from Imperial College London in 2013. Her research focuses on statistical modeling, network analysis, and bioinformatics, with notable contributions to stochastic networks, single-cell genomic data analysis, and complex-valued signal processing. Teaching responsibilities include modules such as Bayesian Methods, Analysing Data, Statistical Analysis, and Project Applied Statistics. She serves as Admissions Tutor for Graduate Certificate and Diploma in Statistics for Data Science and as School Ethics Lead at Birkbeck. Her work bridges theoretical statistics with practical applications in genomics, environmental modeling, and biomedical research. Dr. Chandna’s recent research explores topics like covariate-driven network estimation, stochastic modeling of genomic data, and bootstrap techniques in source separation. Her publications reflect interdisciplinary collaboration across statistics, computer science, and life sciences.
Ozgur Yilmaz is a Professor in the Department of Mathematics at the University of British Columbia (UBC). He is the Director of the Pacific Institute for the Mathematical Sciences (PIMS) and has held roles such as Interim Deputy Director at PIMS and Deputy Director at the Banff International Research Station (BIRS). His research focuses on applied harmonic analysis, signal processing, compressed sensing, and seismic signal processing. Education: PhD in Applied and Computational Mathematics from Princeton University (2001), B.Sc. in Mathematics and Electrical Engineering from Boğaziçi University (1997). Research Interests: Mathematical problems in analog-to-digital conversion, blind source separation, sparse approximations, compressed sensing, and their applications in seismic exploration. He has contributed to advancements in sigma-delta quantization, low-rank matrix recovery, and compressed sensing algorithms. Funding: Recipient of NSERC Discovery Grants, UBC Data Science Institute grants, and leadership in collaborative research groups (CRGs) on high-dimensional data analysis and applied harmonic analysis. His work bridges theoretical mathematics with practical applications in signal processing and AI-driven medical imaging. Students and Postdocs: Supervised numerous PhD and MSc students in areas like compressed sensing, seismic data reconstruction, and machine learning. Current advisees include Aaron Berk and Xiaowei Li. Former students hold positions at academic institutions and tech companies. Labs and Collaborations: Affiliated with UBC’s Data Science Institute (DSI), Centre for Artificial Intelligence Decision-making and Action (CAIDA), and the Institute of Applied Mathematics (IAM). Collaborates on projects integrating AI with scientific discovery, such as retinal biomarker identification using deep learning.
Tetsunori Kobayashi is a Professor in the School of Fundamental Science and Engineering at Waseda University, Japan, where he has served since 1997. He is renowned for pioneering research in human–robot interaction, spoken language processing, and multimodal conversational systems, leading to over 230 refereed papers and an h-index of 35 (Google Scholar). Education: 1980 B.Eng. in Electrical Engineering, Waseda University 1982 M.Eng. and 1985 Dr.Eng. from Graduate School of Science and Engineering, Waseda University Research Interests: His work spans intelligent robotics , perceptual information processing , pattern recognition , image and audio processing , and conversational AI . He develops algorithms for real-time dialogue systems, multi-party conversation facilitation robots, and non-autoregressive speech recognition leveraging CTC and pre-trained language models. Recent Publication Trends: Since 2020 his group has advanced non-autoregressive end-to-end ASR (Mask-CTC, Intermpl, BECTRA), noise-robust attention , multi-look-ahead conversational ASR , and neural speaker diarization . They integrate BERT-style pre-training with CTC losses to accelerate inference while maintaining accuracy. Parallel work explores vision-and-language topics such as scene-graph generation, video semantic indexing, and personalized summarization for spoken news delivery. Scientific Awards: IEICE Fellow 2023 – for multi-modal multi-party conversation research IPSJ Fellow 2016 – for pioneering robot conversation studies JST Award for Academic Start-ups 2024 Best Paper Awards from IEICE, IEEE BTAS, ACM SIGGRAPH VRCAI, and several IPSJ workshop prizes Advising & Grants: He has mentored dozens of PhD and Master’s students who now lead in academia and industry. Major funded projects include JST CREST on conversational robotics, NEDO and JST-support for AI-based speech interfaces, and industry collaborations with NHK, OKI, and NEC. Labs & Teams: Kobayashi heads the Perceptual Computing Laboratory at Waseda, conducting interdisciplinary research with domestic and international partners such as MIT, ATR, and NHK Science & Technology Labs.
Archontis Politis is an Assistant Professor in the Department of Computing Sciences at Tampere University's Faculty of Information Technology and Communication Sciences. His research focuses on signal processing, machine learning, and their applications in audio engineering, particularly in spatial audio, sound source separation, and parametric audio coding. He explores topics such as Ambisonics, reverberation control, and neural network-based approaches for audio processing. His work emphasizes spatial audio reproduction, including six degrees of freedom (6DOF) rendering, microphone array processing, and efficient compression techniques for higher-order Ambisonics. He also investigates sound event localization and detection, leveraging machine learning for real-world acoustic scenarios. His contributions span theoretical advancements in spherical harmonics and practical implementations of spatial audio systems. Recent research highlights include developing datasets for music source separation, improving synthetic-to-real generalization in classical music, and creating neural encoding models for irregular microphone arrays. His methodologies often integrate deep learning with traditional signal processing to address challenges in multi-speaker environments and dynamic acoustic scenes.
Professor Jon Barker is a faculty member at the University of Sheffield , where he holds a Personal Chair in the School of Computer Science . He leads the Speech and Hearing (SpandH) research group and co-founded the CHiME international workshop series on robust speech recognition. Education : PhD in Computer Science (University of Sheffield, 1999); BA in Electrical and Information Sciences (Cambridge University). Research Focus : His work bridges machine listening and human auditory perception , with key contributions to noise-robust speech recognition , speech intelligibility prediction , and hearing aid signal processing for speech and music. Recent projects include the Clarity Challenges and Cadenza Challenges , large-scale machine learning initiatives to improve accessibility for hearing-impaired users. Publication Trends : Recent articles emphasize machine learning for hearing aid optimization , dysarthric speech recognition , audio-visual integration , and music demixing algorithms . Collaborations span speech processing, psychoacoustics, and biomedical engineering. Scientific Awards : EURASIP Best Paper Award (2009) ISCA Best Paper Award (2008) Grants and Leadership : He has secured major EPSRC grants including EnhanceMusic (2022-2026) and Challenges to Revolutionise Hearing Device Processing (2019-2025). He co-led the TAPAS Marie Curie Training Network (2017-2022) and led projects like AV-COGHEAR (2015-2018) and CHiME (2009-2012). Labs and Teams : Barker collaborates closely with the Speech and Hearing Research Group and contributes to international initiatives like the CHiME Workshop . His lab develops open datasets such as the Clarity Speech Corpus and Audio-Visual Lombard Corpus .
Professor Guy Brown is Chair of Computer Science at the University of Sheffield's School of Computer Science. He holds a BSc in Applied Science (1984), PhD in Computer Science (1992), and MEd in Teaching and Learning (1997). His research focuses on Computational Auditory Scene Analysis (CASA), noise-robust speech recognition, auditory modeling, and binaural processing. Research interests include: Machine hearing systems for sound source separation Reverberation-robust speech processing Auditory scene analysis models for normal/impaired hearing Applications in robotics and healthcare technologies Publication trends show recent focus on deep learning approaches for biomedical applications including sleep apnea detection, respiratory sound analysis, and multimodal health monitoring systems using neural networks. Honors include: University Senate Award for Excellence in Teaching (2014) Microsoft Software Engineering Innovation Award (2013) He leads doctoral supervision for 15+ students and has secured research funding from EPSRC, Innovate UK, EU FP7, and AHRC. Manages the Speech and Hearing research group and has held visiting positions at international institutions including LIMSI-CNRS and ATR Japan.
Chenliang Xu is an Associate Professor in the Department of Computer Science at the University of Rochester, affiliated with the Goergen Institute for Data Science and Artificial Intelligence (GIDS-AI). His research focuses on computer vision, audio-visual learning, and trustworthy AI. He holds a PhD from the University of Michigan (2016), with prior degrees from Nanjing University of Aeronautics and Astronautics and the University at Buffalo. Notable awards include the Best Paper Award at ACCV 2024 and the James P. Wilmot Distinguished Professorship. His work spans interdisciplinary topics such as video understanding, multimodal reasoning, and robust AI. Key research contributions include audio-visual scene synthesis, bias mitigation in models, and applications in public health. He has secured over $3M in grants, including NIH funding for AI-driven video description tools and public health initiatives. Prof. Xu advises a dynamic research group with 11 PhD students and numerous collaborators. His lab explores cutting-edge projects like egocentric audio-visual understanding, generative AI for avatars, and multimodal defense mechanisms. He teaches courses in machine vision, deep learning, and advanced computer vision.
Xavier Alameda-Pineda is a Research Director at Inria Grenoble Rhône-Alpes, where he leads the RobotLearn Team. He is affiliated with Université Grenoble Alpes and has been a key member of the Perception team. His work integrates machine learning, computer vision, and audio processing for scene understanding and human-robot interaction. Research Interests: His research lies at the intersection of multimodal machine learning and social behavior analysis. He focuses on developing algorithms for understanding human behavior in natural settings using audio-visual signals, with applications in robotics and AI companions. His work emphasizes real-world challenges such as noisy data, missing modalities, and dynamic environments. Publication Trends: His recent publications reflect a consistent focus on multimodal fusion, particularly combining vision and audio for social scene analysis. Themes include group behavior recognition, sound source separation, and cross-modal learning, often applied in robotics contexts. Scientific Awards: SIGMM Rising Star Award 2018 IEEE TMM Outstanding Associate Editor Award 2022 ACM TOMM Best Paper Award 2020 Best Paper Award, ACM MM 2015 Best Scientific Paper Award, ICPR 2016 Best Student Paper Award, IEEE WASPAA 2015 Outstanding Paper Award, ICMI 2011 Novel Technology Paper Award Finalist, IROS 2017 Advising and Grants: Xavier has mentored students and early-career researchers, evidenced by co-authored student papers. He coordinated the H2020 SPRING project on socially pertinent robots in gerontological healthcare and co-leads an AI chair on audio-visual perception for companion robots, indicating leadership in funded research initiatives. Labs and Teams: He is the leader of the RobotLearn Team at Inria and was previously part of the Perception team. He has also collaborated with the Multimodal and Human Understanding Group at the University of Trento.
Herbert Buchner is a researcher affiliated with the University of Cambridge in the Information Engineering Division , focusing on Machine Learning for Signal Processing and Human-Machine Interfaces . Research Interests : Acoustic scene analysis, biomedical interfaces, haptic systems, wave-domain adaptive filtering, and sensor networks. Applications : Speech recognition, wavefield synthesis, active noise control, and full-duplex communication systems. His work explores TRINICON (a framework for broadband adaptive MIMO filtering), blind source separation, and wave-domain filtering, emphasizing theoretical rigor and real-time implementation. Key Awards : Best Paper Award at ITG Conference on Speech Communication (2008) Best Student Paper Award at IEEE Intl. Workshop on Acoustic Echo and Noise Control (2001) Publications highlight 15 recent articles in areas like: Wave-Domain Adaptive Filtering Blind Source Separation for Convolutive Mixtures Robust Extended Multidelay Filters Multichannel Acoustic Echo Cancellation Active Room Compensation Biomedical Signal Processing
Konrad Kowalczyk is an Associate Professor at AGH University of Science and Technology in Krakow, Poland, where he heads the Signal Processing Group within the Faculty of Computer Science, Electronics and Telecommunications. With extensive international experience from institutions including Queen's University Belfast, Stanford University, and Fraunhofer Institute, he has established himself as a leading researcher in audio and speech signal processing. His academic journey includes B.Eng. and M.Sc. degrees from AGH University (2005), a Ph.D. from Queen's University Belfast (2009), and a Habilitation in ICT from AGH University (2020). B.Eng. and M.Sc. in Electronics and Telecommunications, AGH University of Krakow (2005) Ph.D. in Electronics, Queen's University Belfast, UK (2009) Habilitation (D.Sc.) in Information and Communication Technology, AGH University of Krakow (2020) Kowalczyk's research spans multiple cutting-edge areas in audio processing, with particular focus on speech and audio signal processing enhanced by machine learning techniques. His work integrates deep neural networks with traditional signal processing methods to address challenges in array signal processing , speech enhancement , and speaker recognition . The research group he leads explores innovative applications in distributed signal processing for IoT , acoustic event detection , and spatial audio rendering , bridging theoretical advances with practical implementations. His recent publications demonstrate a clear trend toward integrating deep learning with traditional signal processing techniques, particularly in speaker diarization, source separation, and robust speech recognition. The research increasingly focuses on real-world applications requiring reverberation-robust processing , distributed microphone array systems , and end-to-end neural architectures that can operate in challenging acoustic environments. There's a noticeable shift toward more complex, integrated systems that combine multiple signal processing tasks. Stanislaw Staszic Medal for best graduate of AGH (2005) IEEE Best Student Paper Contest finalist (2007) AES Student Technical Paper Award winner (2008) Best Student Paper Award at IWAENC conference (2014) Best Paper Awards at IEEE SPA conferences (2016, 2019) Polish Ministry of Science Scholarship for Distinguished Young Scientists (2016-2019) Prime Minister Award for outstanding scientific achievements (2020) As Principal Investigator, Kowalczyk leads multiple significant research projects including "Acoustic Intelligence" (2024-2028) funded by National Science Center, and "Deep extraction for robust speech recognition" (2023-2028). He has successfully secured funding from prestigious programs including First TEAM from the Foundation for Polish Science, and EU FP7 projects. His research group actively supervises Ph.D., M.Sc., and B.Eng. students, with strong connections to international institutions including Aalto University and IEEE Signal Processing Society. The research output includes numerous journal publications, conference papers, patents, and software implementations that have advanced the field of audio signal processing. Kowalczyk leads the Signal Processing Group at AGH University, which focuses on developing innovative solutions for speech and audio processing challenges. The group maintains strong collaborations with international institutions including Aalto University (Finland), and participates in European research initiatives. Their work spans theoretical development through practical implementation, with applications ranging from medical voice assistants to distributed acoustic sensor networks.
Mark Plumbley is a Professor of Signal Processing at the Centre for Vision, Speech and Signal Processing (CVSSP) within the School of Computer Science and Electronic Engineering at the University of Surrey. He holds an EPSRC Fellowship in 'AI for Sound' and has led major research initiatives, including the DCASE challenges. His work focuses on AI-driven analysis of acoustic scenes and events, with contributions to machine learning, audio source separation, and sparse representations. Previously, he was Director of the Centre for Digital Music at Queen Mary University of London and Head of the School of Computer Science at Surrey. Education: PhD in Neural Networks (1991). Academic roles include Professorships at King’s College London (1991–2002) and Queen Mary University of London (2002–2014). Research spans audio event detection, sound scene classification, and generative AI for audio synthesis. He leads projects like the EPSRC-funded 'Making Sense of Sounds' and 'Musical Audio Repurposing using Source Separation', and co-edited the Springer book on Computational Analysis of Sound Scenes and Events. Research Interests: AI for Sound: Machine learning applied to real-world audio analysis. Acoustic Scene and Event Recognition: Developing models for sound classification and localization. Generative Audio Models: Text-to-audio systems and diffusion models for sound synthesis. Healthcare Applications: Audio-based diagnostics and bioacoustic signal processing. Grants and Awards: EPSRC Fellowships, EU-funded networks (SpaRTaN, MacSeNet), and Fellowships from IET and IEEE. Notable awards include the IEEE Young Author Best Paper Award (co-authored with students) and leadership in the DCASE community. Labs and Collaborations: CVSSP at Surrey, collaborations with BBC R&D, and interdisciplinary projects on urban soundscapes and noise pollution (UK Acoustics Network Plus).
Slim Essid is a Full Professor at Télécom Paris, leading the Audio Data Analysis and Signal Processing (ADASP) group. He holds a Doctorat (Ph.D.) and Habilitation from Université Pierre et Marie Curie (UPMC). With 15+ years of research experience, he has advised 15 PhD graduates and currently co-advises 10 others. His work focuses on machine learning, signal processing, and multimodal systems, publishing over 150 peer-reviewed papers. He serves as a reviewer for top journals/conferences (e.g., IEEE Transactions) and research funding agencies. Education: State Engineering Degree, École Nationale d’Ingénieurs de Tunis (2001) M.Sc. (D.E.A.) in Digital Communication Systems, École Nationale Supérieure des Télécommunications, Paris (2002) Ph.D., Université Pierre et Marie Curie (2005) Habilitation (HDR), UPMC (2015) Research Interests: Multimodal learning, self-supervised representations, audio-visual segmentation, music structure analysis, domain generalization, and speech enhancement. Recent publications highlight innovations like TACO (training-free sound-prompted segmentation) and CLOUDS (domain-generalized semantic segmentation framework using foundation models). His work bridges audio processing with vision and language models, emphasizing unsupervised/zero-shot approaches. Key achievements include state-of-the-art methods in sound event detection, speaker diarization, and music segmentation. He collaborates with 14 post-docs and leads projects funded by French/EU agencies.