Shinji Watanabe is an Associate Professor at Carnegie Mellon University's Language Technologies Institute and a Courtesy Professor in the Electrical and Computer Engineering department. He holds a Ph.D. (Dr. Eng.) from Waseda University, Japan, and has held research roles at NTT Communication Science Laboratories, Mitsubishi Electric Research Laboratories (MERL), and Johns Hopkins University. His research focuses on automatic speech recognition, speech enhancement, and machine learning for speech processing. Watanabe has published over 300 peer-reviewed papers and received the Best Paper Award at IEEE ASRU 2019. His work emphasizes robust speech processing in challenging environments, multilingual models, and neural audio codecs. He leads the ESPnet toolkit development for end-to-end speech processing systems and contributes to technical committees like IEEE SLTC and APSIPA SLA. Recent research trends include streaming speech systems, universal speech enhancement (URGENT challenges), and fusion of discrete speech units with self-supervised representations. He explores scalable speech foundation models through benchmarks like ML-SUPERB 2.0 and investigates cross-modal audio-visual processing in challenges like MISP 2025. Education : B.S., M.S., Ph.D. (Waseda University) Affiliations : CMU Language Technologies Institute, CMU ECE, Former roles at MERL and Johns Hopkins Key Projects : ESPnet, OpenWhisper-Style Models, URGENT Challenge Frameworks












