James Glass is a Senior Research Scientist at the Massachusetts Institute of Technology (MIT) and heads the Spoken Language Systems Group within MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). He is also affiliated with the Harvard-MIT Division of Health Sciences and Technology. His research spans automatic speech recognition, multimodal learning, and spoken language understanding, with applications in healthcare and video analysis. Education: SM and PhD in Electrical Engineering and Computer Science from MIT His work focuses on paralinguistic speech analysis, health markers in speech, and the intersection of speech and natural language processing. Recent trends emphasize audio-visual alignment, recursive reasoning, and AI applications in cognitive disorder diagnosis. Scientific awards include IEEE Fellow, ISCA Fellow, and Associate Editor for IEEE Transactions on Pattern Analysis and Machine Intelligence. His group explores unsupervised learning, speaker verification, and social text analysis. James leads the Spoken Language Systems Group at CSAIL, collaborating with institutions like IBM and Harvard-MIT Division of Health Sciences and Technology. His research integrates vision-language models, neural audio codecs, and self-supervised frameworks.
Shinji Watanabe is an Associate Professor at Carnegie Mellon University's Language Technologies Institute and a Courtesy Professor in the Electrical and Computer Engineering department. He holds a Ph.D. (Dr. Eng.) from Waseda University, Japan, and has held research roles at NTT Communication Science Laboratories, Mitsubishi Electric Research Laboratories (MERL), and Johns Hopkins University. His research focuses on automatic speech recognition, speech enhancement, and machine learning for speech processing. Watanabe has published over 300 peer-reviewed papers and received the Best Paper Award at IEEE ASRU 2019. His work emphasizes robust speech processing in challenging environments, multilingual models, and neural audio codecs. He leads the ESPnet toolkit development for end-to-end speech processing systems and contributes to technical committees like IEEE SLTC and APSIPA SLA. Recent research trends include streaming speech systems, universal speech enhancement (URGENT challenges), and fusion of discrete speech units with self-supervised representations. He explores scalable speech foundation models through benchmarks like ML-SUPERB 2.0 and investigates cross-modal audio-visual processing in challenges like MISP 2025. Education : B.S., M.S., Ph.D. (Waseda University) Affiliations : CMU Language Technologies Institute, CMU ECE, Former roles at MERL and Johns Hopkins Key Projects : ESPnet, OpenWhisper-Style Models, URGENT Challenge Frameworks
Dr. Andrew Hines is a Researcher at the School of Computer Science, University College Dublin, specializing in machine learning applications for signal processing in speech, audio, and video domains. His work focuses on Quality of Experience (QoE) modeling, speech quality assessment, and immersive media analysis. He has held leadership roles in European COST Actions like Qualinet and CryptoAction, and previously worked in industry as a Director of Engineering. University: University College Dublin Role: Director of Research, Innovation and Impact Key Collaborations: IEEE (Senior Member), Audio Engineering Society (Ireland) Research interests center on machine learning for QoE optimization, audio-visual integration, and healthcare applications like heart sound classification and stroke rehabilitation. His recent publications explore self-supervised learning, neural speech codecs, and contextual factors in speech/audio quality assessment. Scientific contributions include awards like IEEE Senior Membership, and his work spans both academic research and industrial engineering in finance and aviation sectors. He leads the QxLab research team at UCD and develops open-source platforms such as WARP-Q and AQP for quality metrics.
Bryan Pardo is a Professor of Computer Science at Northwestern University and head of the Interactive Audio Lab. He co-directs the Northwestern Center for Human Computer Interaction + Design and chairs the Computer Science Diversity Committee. He teaches courses in Deep Learning, Machine Learning, Generative Modeling, and Digital Music Instrument Design. PhD in Computer Science and Engineering, University of Michigan MMus in Jazz and Improvisation, University of Michigan MS in Computer Science, Ohio State University BMus in Jazz Composition, Ohio State University His research focuses on machine understanding and manipulation of sound, particularly in music and speech domains. Key areas include Machine Learning (e.g., automated gradient clipping), Signal Processing (e.g., Multi-scale Common-fate Transform), and Human Computer Interaction. Applications involve inclusive audio interfaces, audio search engines, source separation, natural language-controlled audio effects, privacy-preserving adversarial attacks on voice recognition, and music co-creation tools. Recent publications highlight advancements in neural watermarking (MaskMark), masked acoustic modeling (VampNet), and real-time adversarial privacy systems for speech. His lab's work has been applied in Adobe's AI-powered audio editor and Lexie B2 hearing aids. Scientific Awards: $1.8 million NSF Future of Work award $440K NSF grant for accessible music programming $200K Toyota grant $100K Sony grant TorchCrepe pitch tracker: 20 million+ downloads Bryan Pardo advises PhD student Max Morrison and collaborates with researchers like Patrick O'Reilly, Zeyu Jin, and Prem Seetharaman. His lab develops technologies for blind and visually impaired audio creators, including HaptEQ and Eyes-free tools.
Farhad Javanmardi is a Researcher at Aalto University within the Department of Information and Communications Engineering. His work focuses on applying advanced computational methods to speech and biomedical signal analysis. Research Interests: Speech processing, voice disorders, machine learning, deep learning, biomedical signal processing, computational linguistics, and health informatics. His recent research trends emphasize the use of transformer-based models and wav2vec2 for robust detection of heart failure and voice pathologies in telephony environments. He investigates database-independent approaches, severity classification, and data augmentation techniques to improve model generalizability. Publications appear in journals like Speech Communication and Computer Speech and Language , as well as conferences including ICASSP and INTERSPEECH .
Chi-Chun Lee (Jeremy) is a Professor and Associate Chair in the Department of Electrical Engineering at National Tsing Hua University (NTHU), Taiwan. He also serves as Director of the NVIDIA-NTHU Joint Innovation Center and leads the Behavioral Informatics & Interaction Computation (BIIC) Lab. His academic journey includes a B.S. (magna cum laude) and Ph.D. in Electrical Engineering from the University of Southern California (USC), USA (2007 and 2012), followed by roles as a data scientist at id:a lab and technical consultant for companies like E.Sun Bank and Allianz Taiwan. Research focuses on speech processing, affective computing, health analytics, and behavior signal processing. He is an IEEE Senior Member and holds editorial roles in top journals such as IEEE Transactions on Affective Computing and Multimedia. Key contributions include leading teams to international competitions (e.g., 1st place in INTERSPEECH 2009 Emotion Challenge) and developing AI frameworks for clinical applications like respiratory sound classification and tumor image synthesis. Recipient of prestigious awards including the NTHU-Novatek Distinguished Talent Chair (2024), National Science and Technology Council Outstanding Research Award (2023), and multiple best paper awards. His work bridges academia and industry, with collaborations extending to NVIDIA and startups like AHEAD Medicine. Research has been featured in major media outlets including Scientific American and Discovery.
Marina Bosi is an Adjunct Professor at Stanford University's Department of Music and a prominent figure in audio engineering and standardization. Currently serving as Chief Technology Officer at MPEG LA, LLC, she has previously held leadership roles including Past-President of the Audio Engineering Society (AES). Education: Laurea (Doctorate) in Physics, University of Florence, Italy Diploma in Flute, National Conservatory of Florence, Italy Honor Diploma in Flute, Accademia Chigiana, Siena, Italy Thesis work with Giuseppe di Giugno at IRCAM, Paris, France Dr. Bosi's research focuses on perceptual audio coding, multichannel audio processing, and signal processing for audio applications. She has been instrumental in developing AI-based media coding standards and audio preservation technologies (ARP), with recent work exploring networked music performance over satellite networks and cross-continental remote collaboration systems. Her publications span foundational audio coding topics including: Psychoacoustic modeling and quantization techniques MDCT/PQMF filter bank design Dolby AC-3 and DAB/DVB multichannel audio coding Bit allocation strategies and quality measurement methods Scientific Awards: AES Board of Governors Award AES Fellowship Award ISO/IEC Special Contribution Award for MPEG-2 AAC development Accademia Nazionale dei Lincei recognition Bourse du Gouvernement Français As a patent holder and author of the seminal textbook Introduction to Digital Audio Coding and Standards , she has significantly shaped modern audio compression technologies. Her professional leadership extends to participation in standards organizations like ANSI, ASA, ATSC, DVB, DVD, ISO, IEC, IEEE, ITU, and SMPTE.
Albert Qiaochu Jiang serves as a Visiting Research Fellow at the Department of Computer Science and Technology, University of Cambridge. His research integrates machine learning with formal theorem proving, focusing on neural theorem provers and mathematical reasoning systems. He leads the reasoning team at Mistral AI while maintaining academic supervision at Cambridge. His research interests center on machine learning for theorem proving , with specific expertise in neural-symbolic integration, autoformalization, and large language models for mathematical reasoning. His work bridges artificial intelligence with formal verification, developing systems that enhance automated reasoning capabilities through neural networks. Current projects involve improving premise selection for theorem provers, multilingual mathematical formalization, and creating efficient architectures for mathematical language models. Analysis of his recent publications reveals a strong focus on advancing neural theorem proving through innovative architectures like Target-Based Automated Conjecturing and Magistral. His research trajectory shows increasing sophistication in integrating language models with formal verification systems, with significant contributions to datasets like Numinamath and frameworks like Llemma. Key trends include optimizing compute efficiency in proof generation, enhancing multilingual mathematical reasoning, and developing interactive human-AI collaboration systems for formal mathematics. While no scientific awards are currently documented in available sources, his research output demonstrates significant impact in the intersection of AI and formal methods. As leader of Mistral AI's reasoning team, Jiang directs research on neural theorem proving systems while contributing to academic supervision at Cambridge. His work involves substantial industrial-academic collaboration, leveraging resources from both institutional contexts to advance mathematical AI. Current projects focus on creating practical systems for mathematical automation with real-world verification applications. His research operates at the intersection of academia and industry through Mistral AI's reasoning team, where he develops neural theorem proving systems with practical applications in formal verification. This dual affiliation enables rapid translation of theoretical advances into deployable tools for mathematical automation.
Antonio Servetti is an Assistant Professor at the Department of Control and Computer Engineering (DAUIN) at Politecnico di Torino, Italy, where he has been a faculty member since 2007. He is affiliated with the Internet Media Group (IMG) and the Interdepartmental Center PIC4SeR for Service Robotics. His work bridges multimedia processing, network communications, and web technologies. MS in Computer Engineering, Politecnico di Torino, 1999 PhD in Computer Engineering, Politecnico di Torino, 2004 Visiting Scholar, University of California, Santa Barbara, 2003 His research focuses on speech and audio processing , multimedia communications over wired and wireless networks , and real-time web-based multimedia applications . Key interests include WebRTC, Web Audio, HTTP adaptive streaming, and perceptual quality assessment. He has contributed to the development of secure multimedia transmission techniques, including selective encryption of speech and audio. The recent publications highlight a strong trend toward AI-driven modeling of subjective quality in multimedia, especially through deep learning for image and video quality prediction, understanding observer behavior, and remote music performance systems. His work often involves collaboration with researchers in the VQEG JEG-Hybrid group and the NEXA Center. Best Paper Award, Web Audio Conference 2021 Dr. Servetti has led and contributed to several research projects, including BRIC-2024 (acoustics in educational settings), PNRR HiFiReM (remote music education), and INAR (artistic research). He teaches courses such as 'Web Applications', 'Machine Learning for Vision and Multimedia', and 'Digital Audio Processing' across various engineering programs. He is also involved in educational governance as a member of academic councils for multiple degree programs. He is a core member of the Internet Media Group (IMG) , which focuses on multimedia processing and transmission, and contributes to the VQEG JEG-Hybrid working group on video quality assessment, where he develops frameworks for reproducible research and modeling of human perception.
Dr. Chang Y Choo is a Professor of Electrical Engineering at San José State University, where he also serves as Director of the AI/ML FPGA/DSP Systems Laboratory. His academic career spans over three decades, with previous positions at Worcester Polytechnic Institute and industry experience at Altera Corp. (now Intel). Dr. Choo maintains an active research program focusing on hardware acceleration for AI and signal processing applications, with particular emphasis on FPGA-based implementations for real-world systems. Dr. Choo's educational background includes: Ph.D. in Computer and Systems Engineering, Rensselaer Polytechnic Institute (1986) M.S. in Operations Research and Statistics, Rensselaer Polytechnic Institute (1982) B.S./M.S. in Engineering, Seoul National University, Korea Dr. Choo's research interests center on the intersection of hardware design and artificial intelligence. His work focuses on implementing computer vision, deep learning, and digital signal processing algorithms on specialized hardware platforms including FPGAs, GPUs, and custom ASICs. Current projects include developing real-time illumination/view-independent object recognition systems for autonomous vehicles, wideband acoustic echo cancellation for wearable technology, and FPGA-based accelerators for medical imaging applications. His research bridges theoretical algorithm development with practical hardware implementation constraints. Analysis of Dr. Choo's recent publications reveals a clear trajectory toward increasingly sophisticated hardware-accelerated AI systems. His work has evolved from foundational research in digital signal processing and image compression to cutting-edge applications of deep learning on specialized hardware. Recent publications demonstrate expertise in implementing CNN architectures on FPGAs, developing metabolic syndrome prediction models, and creating food object detection systems using transformer models. This progression reflects the broader field's shift toward hardware-aware AI development. Dr. Choo's significant scientific contributions include multiple patents that have advanced the state of the art in several domains: U.S. Patent No. 9,025,763 (2015): 'Apparatus and Method for cancelling wideband acoustic echo' U.S. Patent Nos. 7,058,675 (2006) and 7,124,161 (2006): 'Apparatus and method for implementing efficient arithmetic circuits in programmable logic devices' U.S. Patent Nos. 5,943,096 (1999) and 6,621,864 (2003): 'Motion vector based frame insertion process' U.S. Patent Nos. 5,832,131 (1998) and 5,991,455 (1999): 'Hashing-based vector quantization' U.S. Patent No. 5,587,710 (1997): 'Syntax based arithmetic coder and decoder' Throughout his career, Dr. Choo has been actively involved in both academic and industry collaborations. He has served as a technical consultant for numerous Silicon Valley companies including National Semiconductor (now Texas Instruments), Philips Semiconductor, Skybox Imaging (acquired by Google), Novariant (now AgJunction), and Ricoh Innovations. His industry experience informs his teaching approach, which emphasizes practical implementation considerations alongside theoretical foundations. Dr. Choo has also served as an expert witness in intellectual property court cases involving audio and video compression algorithms and FPGA hardware. Dr. Choo directs the FPGA/DSP AI/DL Laboratory at San José State University, which focuses on developing hardware-accelerated solutions for real-time AI applications. The lab maintains strong connections with Silicon Valley technology companies and provides students with hands-on experience in cutting-edge hardware design methodologies. Current research directions include autonomous vehicle navigation systems, medical imaging applications, and edge AI deployment strategies.
Prof. Jörn Ostermann is a Full Professor and Head of the Institut für Informationsverarbeitung at Leibniz Universität Hannover since 2003, with prior roles at AT&T Bell Labs and AT&T Labs-Research. He served as Dean of the Faculty of Electrical Engineering and Computer Science (2011–2013) and member of the Senat (since 2020). His research spans video coding, computer vision, machine learning, 3D modeling, and computer-human interfaces , with applications in SAR imaging, predictive maintenance, children's speech analysis, and cochlear implants. Key projects include Next Generation Video Coding , Conditional Coding for Learned Compression , and GreenAutoML4FAS . Notable trends in his recent publications (2025–2023) include Neural network-based video compression Uncertainty estimation in speech recognition Zero-delay coding for cochlear implants Domain adaptation for aerial image segmentation 3D mesh compression standards Error concealment in VVC coding Scientific recognitions: AT&T Standards Recognition Award (1998) ISO Award (1998) IEEE Fellow (2005) Distinguished Lecturer, IEEE CAS Society (2002/2003) MPEG Convenor (2020–2023) He co-authored a graduate textbook on Video Communications , holds >30 patents, and has led >20 research projects. His work bridges academic research and industrial standardization, particularly in MPEG and IEEE committees.
Professor Kenny Mitchell is a faculty member at the School of Computing Engineering and the Built Environment at Edinburgh Napier University. With over 60 research outputs listed, he specializes in Interactive Graphics, Virtual Reality, and Augmented Reality technologies. Research Interests Mitchell's work focuses on real-time systems, motion prediction, and human-computer interaction in immersive environments. His research spans generative AI environments , 3D facial reconstruction , and light field rendering . Key themes include AI-driven animation , networked VR experiences , and haptic-visual integration . Article Trends Recent publications emphasize Transformer-based motion prediction (NeFT-Net), speech-to-VR systems (HoloJig), and low-latency avatar synchronization . His work integrates machine learning with computer graphics for applications in telepresence dance and emotionally intelligent avatars . Projects CAROUSEL+ : £929,077 funded by European Commission (2021-2024) for telepresent dance systems DISTRO : £243,804 European Commission grant for 3D graphics training (2015-2018)
Lauri Juvela is an Assistant Professor at Aalto University's Department of Information and Communications Engineering, specializing in speech synthesis, audio signal processing, and neural audio effects. His work bridges deep learning, signal processing, and audio engineering. Education: Not explicitly mentioned in provided texts Research Interests include: Speech waveform generation using source-filter vocoding Adversarial speech synthesis and watermarking Virtual analog audio effect modeling with neural networks Nonlinear distortion estimation and restoration Speaker-independent formant synthesis Generative models for audio (GANs, diffusion, DDSP) Recent Article Trends (2025-2021) show focus on: Neural audio effect modeling with synthetic data frameworks Diffusion-based approaches for distortion restoration Nonlinear dynamics linearization in audio effects Collaborative watermarking against adversarial speech High-fidelity glottal excitation models Guitar amplifier modeling with unpaired data Scientific Awards : ISCA Award for best student paper at Interspeech 2016 IEEE Award for best student paper at ICASSP’16 Advising & Grants : No student names mentioned in provided texts. Specific grants not detailed, but active in funded research areas like adversarial speech detection. Labs & Teams : Collaborates with speech synthesis research groups at Aalto University, particularly in the Department of Signal Processing and Acoustics. Works with cross-disciplinary teams in audio codec augmentation and neural audio effect modeling.
Hyunkook Lee is a Professor of Audio and Psychoacoustic Engineering at the University of Huddersfield , where he directs the Applied Psychoacoustics Lab (APL) and the Centre for Audio and Psychoacoustic Engineering . He joined the university in 2010 and has pioneered research in 3D audio, spatial recording, and psychoacoustic applications in virtual reality. Education: PhD in Psychoacoustics (University of Surrey, 2006) Industry Experience: Senior Research Engineer at LG Electronics Research Interests include: Psychoacoustics : Perception of height in 3D audio and interchannel crosstalk. Spatial Audio : Development of microphone arrays, panning, and upmixing techniques. Virtual Reality : Bridging XR and audio engineering for immersive experiences. Music Technology : Applications in commercial rap production and surround sound. Scientific Awards : 5 US-granted patents h-index of 14 (Scopus) Advising & Grants : Actively supervises PhD students and participates in industry-funded projects with Samsung, Volvo, and MagicBeans. His research includes a £326k grant for 3D sound in the 'Audience of the Future'.
Professor Xiang Zhang serves as President and Vice-Chancellor of The University of Hong Kong , combining academic leadership with active research in speech processing and machine learning. His work bridges theoretical and applied advancements in speech enhancement, depression detection, and multilingual speech recognition. Speech Enhancement : Comparative studies on long-context networks and selective state space models Acoustic Representation : Feature codecs, landmark extraction, and frame-rate analysis Mental Health Applications : Depression detection using large language models and speech timing Code-Switching Recognition : Interactive language biasing and alignment techniques Interdisciplinary Research : Unidirectional brain-computer interfaces and quantum language models Recent publications highlight trends in transformer alternatives like Mamba architectures, self-supervised learning for speech tasks, and efficient diffusion-based audio processing. His work spans technical reporting, open-source dataset development, and pedagogical innovations in computational linguistics.