Taiji SuzukiView profile
Associate Professor
Prof. Taiji Suzuki is an Associate Professor at the University of Tokyo in the Department of Mathematical Informatics . He also serves as Team Leader of the "Deep Learning Theory" team at AIP-RIKEN , Japan. With a PhD in Information Science and Technology from the University of Tokyo (2009), he has held academic positions at the University of Tokyo (2009-2013) and Tokyo Institute of Technology (2013-2017) before returning to the University of Tokyo in 2017. University of Tokyo (2004: BEng in Mathematical Engineering) University of Tokyo (2006: MSc in Information Science and Technology) University of Tokyo (2009: PhD in Information Science and Technology) His research interests span statistical learning theory , deep learning , kernel methods , sparse estimation , and stochastic optimization . He investigates how neural networks adapt to function smoothness, avoid the curse of dimensionality, and achieve global optimization through mean-field dynamics. His work bridges theoretical guarantees (minimax optimality, convergence analysis) with practical implementations (transformers, graph neural networks, diffusion models). The 15 most recent articles focus on transformers' representation power, graph neural networks' limitations, diffusion models' convergence, and optimization theories for deep learning. Key themes include information-theoretic bounds , feature learning dynamics , and mean-field analysis . He explores applications in AI for medicine, federated learning, and biomedical modeling. Awards & Recognition: Outstanding Paper Award, ICLR 2021 MEXT Young Scientists’ Prize Outstanding Achievement Award, Japan Statistical Society 2017 Outstanding Achievement Award, Japan Society for Industrial and Applied Mathematics 2016 Best Paper Award, IBISML 2012 Best Paper Candidate, ICDM 2019 He has served as Area Chair for NeurIPS, ICML, ICLR, AISTATS, and as Program Chair for ACML. His scientific contributions include convergence theories for stochastic gradient methods, minimax analysis of deep learning vs kernel methods, and novel frameworks for distributional optimization in diffusion models.







