
معرفی
Anurag Kumar is a Senior Staff Research Scientist at Google DeepMind, focusing on Audio, Speech, and Multimodal AI. His work emphasizes weakly supervised, self-supervised, and unsupervised learning methods. Before DeepMind, he spent six years at Meta and completed his PhD at Carnegie Mellon University (CMU) under Prof. Bhiksha Raj, where his thesis introduced weakly labeled learning of sounds. He holds an Electrical Engineering degree from IIT Kanpur (2013).
Education:
- PhD in Computer Science, Carnegie Mellon University (2013–2018)
- Bachelor of Technology in Electrical Engineering, Indian Institute of Technology (IIT) Kanpur (2008–2013)
Research Interests: Anurag’s research spans audio and speech processing, including speech enhancement, multimodal understanding, and generative AI. He explores techniques like neural radiance fields (NeRF) for audio-visual scene synthesis and diffusion models for music editing. His work bridges theoretical advancements with practical applications in real-world scenarios like room acoustics and egocentric object localization.
Awards & Roles:
- MIT Technology Review 'Innovators Under 35' (2024)
- Associate Editor, IEEE Signal Processing Letters
- Technical Committee Member, IEEE AASP
- Organized workshops on Generative AI for Audio at NeurIPS 2024 and URGENT Challenge for Speech Enhancement
Labs & Collaborations: Leads projects in multimodal AI at Google DeepMind, collaborating on tools like Torchaudio and frameworks for audiovisual learning. His work includes benchmark datasets like Real Acoustic Fields and advancements in neural field-based scene synthesis.


