Dr. Chang Xu is an Associate Professor in Machine Learning and Computer Vision at the University of Sydney's School of Computer Science. He holds a Bachelor of Engineering from Tianjin University and a PhD from Peking University. His research focuses on machine learning, data mining, and their applications in AI and computer vision, including multi-view learning, visual search, and face recognition. He is an ARC Future Fellow and a member of the Sydney Southeast Asia Centre and The Net Zero Institute. Education: B.E. in Engineering (Tianjin University), Ph.D. in Computer Science (Peking University). His research interests emphasize handling heterogeneous data, exploring data variety, and developing algorithms for robust AI systems. His work includes adversarial robustness, neural architecture search, and efficient deep learning models. Research trends in his articles include adversarial robustness in neural architectures, efficient vision transformers, multimodal 3D style transfer, and underwater image restoration. Key contributions span image restoration, video super-resolution, and lightweight network design. He has advised multiple PhD and master's students on topics like diffusion models, radar image synthesis, and graph similarity. Awards: ARC Future Fellow. Collaborations focus on cross-domain data integration and AI applications. His labs and teams explore generative models, robust learning, and scalable robotics policies. Recent work includes diffusion models for action segmentation and robust vision-language systems.
Kristen Grauman is a Full Professor in the Department of Computer Science at the University of Texas at Austin, where she leads the UT Computer Vision Group. Her research focuses on computer vision and machine learning, with applications in visual recognition, video analysis, and multi-modal perception. She received her B.A. from Boston College and her Ph.D. from MIT. Her research interests span visual recognition, image and video search, video analysis, first-person vision, embodied and multi-modal perception, and interactive machine learning. She has made significant contributions to the field, particularly in developing algorithms for understanding visual content and human activities from video, including foundational work on the Pyramid Match Kernel and relative attributes. Her recent publications reveal a strong emphasis on egocentric (first-person) vision, audio-visual learning, and view-invariant representations. There is a clear trajectory toward multi-modal integration (vision, audio, language) and real-world applications in instructional videos, human activity understanding, and embodied AI systems. She has received numerous awards including: AAAI Fellow (2019) J. K. Aggarwal Prize, International Association for Pattern Recognition (2018) Helmholtz Prize (2017) UT Austin Academy of Distinguished Teachers (2017) Best Paper Award, Asian Conference on Computer Vision (2016) Presidential Early Career Award for Scientists and Engineers (2014) Computers and Thought Award, International Joint Conferences on Artificial Intelligence (2013) Pattern Analysis and Machine Intelligence Young Researcher Award (2013) Alfred P. Sloan Research Fellow (2012) Marr Prize (2011) Prof. Grauman serves as Associate Editor-in-Chief for the IEEE Transactions on Pattern Analysis and Machine Intelligence. She has secured substantial research funding including the Presidential Early Career Award, NSF grants, and industry partnerships. Her advising has produced numerous influential publications and students who are now leaders in computer vision. She leads the UT Computer Vision Group, which collaborates closely with the Electrical and Computer Engineering Department. The group is pioneering large-scale egocentric video research through projects like Ego4D and Ego-Exo4D, focusing on real-world applications in human activity understanding, audio-visual perception, and interactive systems.
Katerina Fragkiadaki is the JPMorgan Chase Associate Professor of Computer Science in the Machine Learning Department at Carnegie Mellon University. She works at the intersection of Artificial Intelligence, Computer Vision, Machine Learning, Language Understanding, and Robotics. PhD from GRASP Lab, University of Pennsylvania Postdoctoral researcher at UC Berkeley (with Jitendra Malik) and Google Research Recipient of NSF CAREER, DARPA Young Investigator, Amazon, Google, Sony, UPMC, and AFOSR awards Organizer of CoRL 2023 Workshop on Generalist Robots ICLR 2024 Program Chair, multiple area chair roles Her research group focuses on developing machines that autonomously improve world models through human-environment interactions, with specific emphasis on: Representation learning and video understanding 2D/3D unified vision-language models Generative simulation and reinforcement learning Real2Sim/Sim2Real robot learning Continual learning and spatial common sense 3D scene reconstruction and dynamics Recent publications highlight advancements in: 3D mesh generation with compositional transformers Unified 2D/3D perception frameworks Physics-aware generative models Diffusion-based robotic manipulation policies Embodied agents with memory prompting Awards include: 2024: DARPA Young Investigator Award 2023: Amazon Faculty Award 2022: Sony Faculty Research Award 2021: UPMC Faculty Research Award 2020: NSF CAREER Award 2019: Google Faculty Award Key collaborations span institutions including UC Berkeley, Google Research, Stanford, MIT, and University of Tsukuba. Her work bridges theoretical innovation with practical applications in: Autonomous robot manipulation 4D world modeling Language-grounded perception Visual dynamics prediction Embodied program synthesis Physics-based simulation engines
Andrew Zisserman is a Royal Society Research Professor at the University of Oxford's Department of Engineering Science, affiliated with the Visual Geometry Group (VGG). His research focuses on computer vision, artificial intelligence, and neural networks, with significant contributions to multimodal learning, video understanding, and 3D scene analysis. He leads projects exploring visual-language models, audio-visual synchronization, and clinical imaging applications. Key research areas include: Video analysis and temporal modeling Multimodal systems for sign language translation and action recognition 3D shape estimation and physical property inference Foundation models and cross-modal retrieval Recent work highlights: Developed Flamingo and Tapir models for video-language tasks Advancements in spinal MRI analysis and clinical imaging Leadership in EGO4D and VoxCeleb challenges Honors include Fellowship of the Royal Society (FRS) and the ISSLS Prize in Clinical Science 2023 for spinal analysis innovations. His lab collaborates globally, emphasizing real-world applications in healthcare and autonomous systems.
Ajit Rajwade is a Professor at the Department of Computer Science and Engineering , Indian Institute of Technology Bombay. His research spans Artificial Intelligence , Compressed Sensing , and Medical Imaging , with affiliations to the Centre for Machine Intelligence and Data Science and the Koita Centre for Digital Health . Education : PhD in Computer and Information Science and Engineering, University of Florida (2010) MSc in Computer Science, McGill University (2004) BTech in Computer Engineering, University of Pune (2002) His research interests focus on intelligent data acquisition, particularly in neural network analysis , graph signal processing , and medical imaging . He develops compressed sensing algorithms for inverse problems like tomography and MRI reconstruction , alongside applying group testing to pandemic response. Scientific awards include the Prof. S. P. Sukhatme Award (2024) and Departmental Teaching Excellence (2019). His publications demonstrate expertise in image restoration , noise modeling , and epidemiological algorithms . He has advised PhD students like Sabyasachi Ghosh and Jerin Geo James , with a focus on computationally efficient methods in medical imaging and machine learning .
Timothy M. Hospedales is a Professor of Artificial Intelligence at the Institute of Perception, Action and Behaviour within the School of Informatics at the University of Edinburgh . He also serves as VP AI and Head of Samsung AI Research Centre Europe . His research focuses on efficient and robust AI , emphasizing meta-learning , lifelong transfer-learning , and domain adaptation in both probabilistic and deep learning frameworks. Applications span computer vision , vision and language , reinforcement learning for robotics , and finance . Professor at University of Edinburgh (2020–present) ELLIS Fellow (2021) Head of Samsung AI Research Europe (2020–present) Founding Director of Applied Machine Learning Lab at QMUL (2012–2016) His work includes pioneering contributions to meta-learning , few-shot learning , and self-supervised methods , with notable awards such as the Best Paper Prize at ICML AutoML 2018 and Best Student Paper at ICPR 2018 . He has co-authored 15+ recent papers on topics like Vision-Language Models , Medical AI Fairness , and Diffusion Model Optimization . He served as Program Co-Chair for BMVC 2018 and AAAI 2022 , and authored a book on Visual Adaptation in the Deep Learning Era (2022). Co-Chair, BMVC 2018 Guest Editor, IET CV Special Issue (2016) Keynote Speaker at TASK-CV Workshop (ECCV 2016) Special Issue on Fewer Labels (IEEE PAMI 2020) His leadership extends to organizing workshops like the Learning-to-Learn Workshop at ICLR 2021 , Meta-Learning Workshop at NeurIPS 2020 , and Domain Generalisation Workshop at ICLR 2023 . Current projects include Meta-Omnium (CVPR 2023) for general-purpose meta-learning and MetaAudio (ICANN 2022) for few-shot audio classification benchmarks.
Judy Hoffman is an Associate Professor in the College of Computing at Georgia Institute of Technology, with a joint appointment in the School of Interactive Computing and affiliation to the Machine Learning Center . She received tenure in April 2025 after joining Georgia Tech as an Assistant Professor. Her research focuses on enabling AI systems that are reliable, fair, and resource-efficient. PhD in Electrical Engineering and Computer Science (2016, UC Berkeley) Postdoctoral Fellowships at Stanford (2017) and UC Berkeley (2018) Former Research Scientist at Facebook AI Research Her work intersects computer vision and machine learning , with specialization in domain adaptation , adversarial robustness , and algorithmic fairness . She has published over 40 peer-reviewed articles, including the award-winning DeCAF (ICML 2024 Test of Time Award) and co-founded Women in Computer Vision (2015), which has sponsored ~40 women annually to premier conferences. ICML Test of Time Award (2024) NSF CAREER Award (2022) PAMI Distinguished Young Researcher (2023) Samsung AI Researcher of the Year (2021) Dr. Hoffman has served as Program Chair for CVPR 2023, Associate Editor for T-PAMI (2021-2023), and co-organizer of workshops at major AI conferences. She has delivered over 70 invited talks and contributes to open-source projects like cycada_release (567 stars) and lsda (47 stars).
Professor Winston Hsu is a distinguished faculty member in the Department of Computer Science and Information Engineering at National Taiwan University, where he has served as a full professor since 2015. He is the co-director of the Communications and Multimedia Laboratory (CMLab) and founder of the MiRA (Multimedia indexing, Retrieval, and Analysis) research group. Additionally, he serves as the Founding Director for NVIDIA AI Lab at NTU, the first such lab in Asia. Professor Hsu received his Ph.D. in Electrical Engineering from Columbia University in 2007 under the supervision of Professor Shih-Fu Chang. Prior to his academic career, he was a founding engineer and research manager at CyberLink Corp., now a public image/video software company. National Taiwan University (2007-Present): Professor (2015-Present), Assistant/Associate Professor (2007-2015) MobileDrive (2021-2024): CTO and Vice President (Joint Venture between Foxconn and Stellantis) IBM TJ Watson Research Center (2016-2017): Visiting Scientist Microsoft Research Redmond (2014): Visiting Researcher Columbia University (2007): Ph.D. in Electrical Engineering Professor Hsu's research focuses on machine learning, computer vision, large-scale image and video search and recognition, and embedded AI. His work spans from fundamental research in visual recognition to practical applications in automotive systems, medical imaging, and e-commerce. He has pioneered work in disguised face recognition, low-resolution face hallucination, 3D model search, and virtual try-on systems. His current research emphasizes Embodied AI, integrating perception, action, and learning technologies for applications in automotive and robotics domains. His research group has produced numerous influential publications, particularly in top computer vision and multimedia conferences like CVPR, where they won first place in the Disguised Face Recognition competition in 2018. Their work spans diverse application areas including security, medical diagnostics, automotive systems, and e-commerce solutions, demonstrating strong translation from academic research to real-world impact. IBM Research Pat Goldberg Memorial Best Paper Award (2018) First Place, IARPA Disguised Faces in the Wild Competition (CVPR 2018) Best Brave New Idea Paper Award, ACM Multimedia 2017 NVIDIA AI LAB Award (First in Asia, 2016) First Place, MSR-Bing Image Retrieval Challenge (2013) World's Top 2% Scientists (2023) Professor Hsu actively mentors students and researchers, with his group consistently recruiting PhD students, postdocs, and research assistants. He has successfully bridged academia and industry through multiple collaborations, including his role as CTO at MobileDrive (a joint venture between Foxconn and Stellantis) from 2021-2024. His research has been supported by significant industry partnerships with Microsoft, IBM, and NVIDIA, as well as government grants from Taiwan's Ministry of Science and Technology. His laboratory, the Communications and Multimedia Laboratory (CMLab), maintains strong industry connections and focuses on cutting-edge research in visual AI. The lab has produced numerous award-winning projects and maintains active collaborations with global technology companies, particularly in the automotive and consumer electronics sectors.
Yanzhi Wang is a Professor in the Department of Electrical and Computer Engineering at Northeastern University , affiliated with the Institute for Experiential AI and the Institute for the Wireless Internet of Things . He holds a PhD from the University of Southern California (2014). His research focuses on real-time AI systems, deep neural network compression, neuromorphic computing, and non-von Neumann architectures. Notable projects include NSF-funded initiatives on age-inclusive urban design, superconducting computing (DISCoVER), and edge device optimization (PatDNN). He has received prestigious awards such as the Army Research Office Young Investigator Award and the Constantinos Mavroidis Translational Research Award. His work emphasizes algorithm-hardware co-design for energy efficiency, with grants from NSF, ARO, and industry partners like Google. Recent research trends reflect his focus on accelerating vision transformers, diffusion models, and large language models for edge computing. He has pioneered methods like AutoViT and Fastcar, addressing latency and resource constraints in mobile platforms. Collaborations span academia and industry, driving innovations in superconducting circuits and neuromorphic systems.
Bryan A. Plummer is an Assistant Professor in the Department of Computer Science at Boston University, affiliated with the IVC Group and the Artificial Intelligence Research (AIR) initiative at the Rafik B. Hariri Institute. He holds a PhD from the University of Illinois at Urbana-Champaign, specializing in computer vision. His research focuses on multimodal machine learning, efficient neural architectures, explainable AI, and robust ML systems. Plummer's work bridges vision and language, addressing challenges in domain generalization, synthetic data utilization, and model efficiency. Notable contributions include the Flickr30K Entities dataset and advancements in vision-language model robustness against web artifacts. He has advised over 20 students, with several securing roles at top institutions like NVIDIA and Google. His recent awards include the 3M Foundation Fellowship and NSF GRFP honorable mention. Plummer actively serves on conference committees (NeurIPS, CVPR, ICCV) and leads initiatives like the 1st Findings Workshop at ICCV'25.
Ron Fedkiw is the Canon Professor of Computer Science at Stanford University's School of Engineering. He holds a PhD in Applied Mathematics from UCLA. His research focuses on computational algorithms for applications in computational fluid dynamics, computer graphics, biomechanics, and machine learning. Fedkiw has pioneered techniques for simulating natural phenomena in film and video games, earning two Academy Awards for his contributions to visual effects. He leads the PhysBAM lab and collaborates with industry through consulting roles at Epic Games and former work with Industrial Light & Magic. Education: PhD in Applied Mathematics, UCLA (1996). Notable awards include the National Academy of Science Award, Packard Fellowship, and multiple teaching honors. His lab has graduated 40 PhD students, many of whom have made significant impacts in academia and industry. Research interests span fluid dynamics, cloth simulation, facial animation, and integrating machine learning with physical models. Key contributions include algorithms for two-way fluid-solid coupling, muscle-based facial modeling, and neural network approaches for cloth and deformable bodies. Current projects explore physics-informed machine learning and real-time interactive simulations. Scientific Awards include two Oscars, PECASE, and Okawa Foundation grants. His work bridges computational physics and visual effects, with over 140 research papers and a textbook on level set methods. Advising and grants: Supervised 40 PhD students, securing funding through NSF, ONR, and industrial partnerships. Lab collaborations include SAIL (Stanford AI Lab) and Epic Games. Future work focuses on AI-driven physical simulations and biomedical applications.
Wenhu Chen is an Assistant Professor at the University of Waterloo's Computer Science Department and a CIFAR AI Chair at the Vector Institute. He also holds a part-time role as a Senior Research Scientist at Google DeepMind (20% allocation). His research focuses on natural language processing, deep learning, and multimodal reasoning, with contributions to models like MAmmoTH, OpenCoderInterpreter, and VISTA. He received awards including the Canada CIFAR AI Chair (2022) and the UCSB CS Outstanding Dissertation Award (2021). Education: PhD in Computer Science from the University of California, Santa Barbara (under William Wang and Xifeng Yan). Research interests include complex reasoning, controllable GenAI, and multimodal benchmarks like MEGABench and MMMU. Grants include CIFAR AI Chair Funding (2022-2027), NSERC Discovery Fund (2023-2028), and multiple NRC Canada grants. He directs the TIGER Lab, advancing generative models in text, images, videos, and music. Recent talks include presentations on multimodal reasoning at Apple and NeurIPS workshops.
Howie Choset is a Professor of Robotics at the Robotics Institute, Carnegie Mellon University. He directs the Undergraduate Robotics Minor and leads the Biorobotics Laboratory, where his research focuses on snake robots, motion planning, and medical robotics. He is also affiliated with the Manufacturing Futures Institute. Ph.D., Mechanical Engineering, California Institute of Technology (1996) M.S., Mechanical Engineering, California Institute of Technology (1995) B.S.E., Computer Science and Engineering, University of Pennsylvania (1990) B.S., Economics, The Wharton School of Business (1990) Choset's research centers on robotics for confined and complex environments, particularly through the development of snake robots. His work integrates mechanism design, path and motion planning, and estimation to enable applications in surgery, manufacturing, infrastructure inspection, and search and rescue. He is a pioneer in medical robotics and has co-founded Medrobotics to commercialize minimally invasive surgical robots. His recent publications highlight a strong focus on ergodic exploration, multi-agent systems, motion planning under uncertainty, and medical robotics. Themes include optimizing robot trajectories for information gathering, solving complex path planning problems with dynamic obstacles, and advancing autonomous systems for disaster response and space exploration (e.g., the EELS robot for Enceladus). MIT Technology Review Top 100 Innovators under 35 (2002) Best Paper Award, RIA (1999) Best Paper Award, ICRA (2003) Best Paper, IEEE Bio Rob (2006) Best Video, ICRA (2011) Nominations for best papers at ICRA, IROS, and CLAWAR Choset has advised numerous students, many of whom have won top awards. His lab has received significant funding for robotics research, including projects in surgical robotics, additive manufacturing, and autonomous exploration. He is the lead author of the textbook Principles of Robot Motion and is actively involved in educational innovation through custom robotics labs. He leads the Biorobotics Laboratory at CMU, which develops advanced robotic systems like snake robots and the EELS (Exobiology Extant Life Surveyor) robot for NASA missions. The lab collaborates with industry and government agencies on applications ranging from surgery to space exploration.
Shih-Fu Chang is the Dean of Columbia Engineering and holds the Morris A. and Alma Schapiro Professorship at Columbia University. His research focuses on computer vision, machine learning, and multimedia information retrieval. He is recognized as a foundational figure in the field of content-based visual search and has pioneered innovations in image/video search engines, crime prevention systems, and brain-machine interfaces. His leadership roles include Chair of Columbia's Electrical Engineering Department (2007-2010), Editor-in-Chief of the IEEE Signal Processing Magazine (2006-2008), and Senior Executive Vice Dean at Columbia Engineering, where he drives strategic planning and international collaboration. Dr. Chang has received prestigious awards including the ACM Multimedia Technical Achievement Award, IEEE Signal Processing Technical Achievement Award, and IEEE Kiyo Tomiyasu Award. He is a Fellow of AAAS, ACM, and IEEE, and an Academician of Academia Sinica. His recent work emphasizes multimodal reasoning, few-shot learning, and vision-language systems, with applications in healthcare diagnostics and multimedia benchmarking. His research spans cross-modal understanding, event extraction, and adaptive AI systems. Key contributions include systems like Ferret-v2 for multimodal grounding and RESIN for schema-guided event tracking. He has advised multiple startups and actively contributes to curriculum development in AI and engineering education.
Oswald Lanz is a tenured full professor at the Faculty of Engineering of the Free University of Bozen-Bolzano , leading the Visual Computing Lab . He holds a Ph.D. in Computer Science and a Mathematics degree from the University of Trento. Prior to his current role, he was a researcher and head of research at FBK Trento. He is an endowed professor collaborating with Covision Lab , an AI hub in Bressanone, and coordinates the board of professors for the PhD in Computer Science program since 2025. His research focuses on Computer Vision, Deep Learning, and Video Analytics , with applications in sports technology, medical imaging, and industrial automation. Key achievements include the Amazon AWS Machine Learning Research Award (2020) , ACM Multimedia Best Paper (2015) , and Best Student Paper at ICIAP (2007) . He co-organized the ELLIS-VISMAC Winter School (2025) and chaired ICIAP 2019 . His work spans novel view synthesis, action recognition, and anomaly detection, supported by patents in video tracking and detection. He teaches courses like Deep Learning and Artificial Intelligence in undergraduate and graduate programs. Recent projects such as 5VREAL integrate 5G, edge computing, and AI for sports analysis. His collaborations bridge academia and industry, exemplified by his role in Covision Lab and multidisciplinary initiatives like DSS4LCO for food supply chains. Lanz’s publications emphasize spatiotemporal modeling, neural architecture search, and hybrid machine vision systems.