Carl Vondrick is a Professor in the Department of Computer Science at Columbia University. His research focuses on creating robust and versatile perception systems that leverage video and interaction with the natural world, with applications in 3D reconstruction, visual question answering, and robot manipulation. Former research scientist at Google Visiting researcher at Cruise Education: PhD (2017) from MIT, advised by Antonio Torralba BS (2011) from UC Irvine, advised by Deva Ramanan His research explores multimodal approaches for cross-task and cross-modal transfer, scene dynamics, audiovisual perception, interpretable models, and spatial awareness systems. The lab emphasizes zero-shot generalization and neuro-symbolic methods while addressing safety and robustness in AI systems. Key publication trends include: 2025: Video generation for robotics 2024: Differentiable rendering and cross-modal reasoning 2023: Robust perception and 3D modeling Scientific Awards: 2024 PAMI Young Researcher Award 2021 NSF CAREER Award Teaching Roles: Teaching Computer Vision II (2021-2025), Computer Vision I (2018-2019), and Representation Learning (2020-2022). Advising: Advises 8 current PhD students and has mentored 5 graduated students now at institutions like MBZUAI and UMD. The lab recruits 1-2 PhD students annually through Columbia’s PhD program. Grants and Collaborations: Funded by NSF, DARPA, Toyota Research Institute, Amazon Research, and Google.
Prof. Matthias Nießner is a Professor at the Technical University of Munich, leading the Visual Computing Lab. His research intersects computer graphics, vision, and AI, focusing on 3D reconstruction, semantic understanding, and AI-driven video synthesis. He holds a PhD from the University of Erlangen-Nuremberg (2013) and was a Visiting Assistant Professor at Stanford University (2013–2017). Notable awards include the ERC Starting Grant (2018), Nvidia Professorship Award, and Eurographics Young Researcher Award (2019). His work has been featured in mainstream media and led to startups like Synthesia Inc. Research spans Gaussian splatting, neural radiance fields, and generative AI for 3D avatars. Over 150 publications include SIGGRAPH, CVPR, and ECCV, with best paper awards. Projects like Face2Face and ScanNet have driven innovation in facial reenactment and 3D scene datasets. Education: PhD in Computer Science, University of Erlangen-Nuremberg (2013) Diploma in Computer Science, University of Erlangen-Nuremberg (2010) Research Interests: 3D digitization, neural rendering, generative AI, non-rigid reconstruction, and applications in AR/VR. Awards: ERC Starting Grant (2018) Nvidia Professorship Award (2018) Google Faculty Award (2018) SIGGRAPH Best Emerging Tech Award (2016) Grants: Over €1.5M from ERC and industry partnerships. Labs/Teams: Visual Computing Lab at TUM and Synthesia Inc. (co-founder). Key projects include ScanNet (large 3D indoor dataset), Face2Face (real-time facial reenactment), and Gaussian-based 3D avatars. Current work focuses on diffusion models, neural radiance fields, and AI-generated media detection.
Achuta Kadambi, Ph.D., is an Associate Professor at UCLA in Electrical Engineering and Computer Science, leading an interdisciplinary research group focused on AI, computational imaging, and bias mitigation in medical technologies. He recruits PhD students from EE, CS, and Bioengineering departments and has commercialized research through two California-based companies. His research investigates the intersection of physics and artificial intelligence, with a focus on unbiased low-level vision systems. Current projects explore how light transport interacts with human skin variations to identify and correct imaging biases in facial recognition and medical devices. His work has produced over 70 patents, with 30+ issued, and a textbook Computational Imaging (MIT Press, 2022). NSF CAREER Award (2021) for light transport bias research DARPA Young Faculty Award (2021) for AI and medical imaging innovations ARO Young Investigator Program (2021) for computational sensing IEEE-HKN Under 35 Award (2022) for inclusive EECS inventions Forbes 30 Under 30 recognition His recent publications focus on polarization imaging, 3D Gaussian splatting, synthetic data generation for healthcare, and bias mitigation in machine learning. Collaborations with UCLA medical school faculty, including Dr. Laleh Jalilian, aim to deploy these innovations in clinical settings. Current teaching includes ECE 149: Foundations of Computer Vision (Fall 2024, Spring 2025) and ECE 102: Signals and Systems (Winter 2024).
Jens Edlund is a Professor at KTH Royal Institute of Technology's Division of Speech, Music and Hearing. His research focuses on speech technology, dialogue systems, prosody, and evolutionary phonetics. He has contributed to foundational work on speech synthesis, conversational interaction, and multimodal corpora like the D64 corpus. Key projects include the MonAMI Reminder system and analysis of primate vocalizations to understand speech evolution. Edlund has collaborated extensively with global researchers, producing over 150 peer-reviewed works. His work integrates computational methods with linguistic and biological insights, emphasizing human-like dialogue systems and cross-species vocal analysis. Education: Ph.D. in Speech Technology (2011, KTH) Grants: Multiple EU and Swedish Research Council grants for speech technology and interdisciplinary studies Research labs include the KTH Speech, Music and Hearing Lab and collaborations with institutions like Max Planck Institute for Evolutionary Anthropology. Current work explores evolutionary origins of speech biomechanics and AI-driven speech synthesis evaluation.
Prof. Matthias Nießner is a Professor at the Technical University of Munich , where he leads the Visual Computing Lab . Prior to this, he held a Visiting Assistant Professor position at Stanford University . His work bridges computer vision , graphics , and machine learning , focusing on 3D reconstruction , semantic scene understanding , and AI-driven video synthesis . Prof. Nießner has published over 150 works in top venues like SIGGRAPH , CVPR , and ECCV , with several receiving best paper awards (SIGCHI’14, HPG’15, SPG’18, SIGGRAPH’16 Emerging Tech). His research has garnered international media attention, including features in the New York Times , Wall Street Journal , and MIT Technological Review , as well as TV demonstrations (e.g., Jimmy Kimmel Live for Face2Face technology). Awards : TUM-IAS Rudolph Moessbauer Fellowship (2017–ongoing) Google Faculty Award (2017) Nvidia Professor Partnership Award (2018) ERC Starting Grant (2018, €1.5M) Eurographics Young Researcher Award (2019) Research Trends : 3D Gaussian Splatting for real-time rendering Neural Radiance Fields (NeRF) with mesh supervision Audio-driven facial animation via diffusion models Latent space diffusion for 3D scenes Self-supervised and zero-shot methods for 3D and image analysis As a co-founder and director of Synthesia Inc. , he drives democratization of synthetic media. His YouTube channel has over 5 million views, reflecting his impact beyond academia.
Professor Winston Hsu is a distinguished faculty member in the Department of Computer Science and Information Engineering at National Taiwan University, where he has served as a full professor since 2015. He is the co-director of the Communications and Multimedia Laboratory (CMLab) and founder of the MiRA (Multimedia indexing, Retrieval, and Analysis) research group. Additionally, he serves as the Founding Director for NVIDIA AI Lab at NTU, the first such lab in Asia. Professor Hsu received his Ph.D. in Electrical Engineering from Columbia University in 2007 under the supervision of Professor Shih-Fu Chang. Prior to his academic career, he was a founding engineer and research manager at CyberLink Corp., now a public image/video software company. National Taiwan University (2007-Present): Professor (2015-Present), Assistant/Associate Professor (2007-2015) MobileDrive (2021-2024): CTO and Vice President (Joint Venture between Foxconn and Stellantis) IBM TJ Watson Research Center (2016-2017): Visiting Scientist Microsoft Research Redmond (2014): Visiting Researcher Columbia University (2007): Ph.D. in Electrical Engineering Professor Hsu's research focuses on machine learning, computer vision, large-scale image and video search and recognition, and embedded AI. His work spans from fundamental research in visual recognition to practical applications in automotive systems, medical imaging, and e-commerce. He has pioneered work in disguised face recognition, low-resolution face hallucination, 3D model search, and virtual try-on systems. His current research emphasizes Embodied AI, integrating perception, action, and learning technologies for applications in automotive and robotics domains. His research group has produced numerous influential publications, particularly in top computer vision and multimedia conferences like CVPR, where they won first place in the Disguised Face Recognition competition in 2018. Their work spans diverse application areas including security, medical diagnostics, automotive systems, and e-commerce solutions, demonstrating strong translation from academic research to real-world impact. IBM Research Pat Goldberg Memorial Best Paper Award (2018) First Place, IARPA Disguised Faces in the Wild Competition (CVPR 2018) Best Brave New Idea Paper Award, ACM Multimedia 2017 NVIDIA AI LAB Award (First in Asia, 2016) First Place, MSR-Bing Image Retrieval Challenge (2013) World's Top 2% Scientists (2023) Professor Hsu actively mentors students and researchers, with his group consistently recruiting PhD students, postdocs, and research assistants. He has successfully bridged academia and industry through multiple collaborations, including his role as CTO at MobileDrive (a joint venture between Foxconn and Stellantis) from 2021-2024. His research has been supported by significant industry partnerships with Microsoft, IBM, and NVIDIA, as well as government grants from Taiwan's Ministry of Science and Technology. His laboratory, the Communications and Multimedia Laboratory (CMLab), maintains strong industry connections and focuses on cutting-edge research in visual AI. The lab has produced numerous award-winning projects and maintains active collaborations with global technology companies, particularly in the automotive and consumer electronics sectors.
Yaser Sheikh is an Associate Professor at the Robotics Institute of Carnegie Mellon University (on leave) and Director of the Facebook Reality Lab, Pittsburgh . He holds appointments in the Mechanical Engineering Department and focuses on ' metric telepresence ' for AR/VR interactions. His research spans machine perception , computer vision , computer graphics , and machine learning , with applications in social behavior modeling and dynamic 3D reconstruction. University: Carnegie Mellon University Roles: Associate Professor (Robotics Institute), Director (Facebook Reality Lab) Contact: yaser@cs.cmu.edu, yasers@fb.com Research Interests include: Computer Vision: Pose estimation, 3D reconstruction, camera calibration Computer Graphics: Face/Hand animation, photorealistic rendering Machine Learning: Neural rendering, unsupervised learning for landmark detection AR/VR: Telepresence, immersive social interactions Notable Trends in Publications reveal a focus on real-time pose estimation (e.g., OpenPose), dynamic 3D reconstruction , and codec avatars for VR/AR. Recent works emphasize universal priors and neural rendering for photorealistic avatars. Scientific Awards include: Popular Science’s Best of What’s New Award Honda Initiation Award (2010) Best Paper Awards: WACV (2012), SCA (2010), ICCV THEMIS (2009) Hillman Fellowship for Excellence in Computer Science Research (2004) Advising and Grants: He has advised numerous PhD students (e.g., Hanbyul Joo, Tomas Simon) and received funding from the National Science Foundation , DARPA, and industry partners like Intel , Disney , and Honda . Labs & Teams: Leads the Facebook Reality Lab in Pittsburgh, collaborating with institutions like Carnegie Mellon University and Disney Research.
Zakia Hammal is an Assistant Research Professor with dual appointments at Carnegie Mellon University, holding positions in the Robotics Institute within the School of Computer Science and the Department of Biomedical Engineering in the College of Engineering. Her work bridges computer science, machine learning, artificial intelligence, and social/behavioral psychology to advance computational models for human behavior analysis. Dr. Hammal's educational background includes a PhD in Computer Science, a Master of Artificial Intelligence and Algorithmic with specialization in Image Processing, and an Engineer's degree in Computer Science with specialization in Computer Systems. Her academic journey has positioned her at the intersection of technical expertise and healthcare applications. Her research focuses on multimodal human behavior modeling in social interaction, with particular emphasis on health informatics and affective computing (Emotion AI). Dr. Hammal's work has pioneered computational models for multimodal assessment of psychiatric disorders, including depression severity evaluation, automatic pain intensity measurement, assessment of expressiveness in children with facial abnormalities, analysis of non-verbal communication in mother-infant interaction, and identification of behavioral markers in autism spectrum disorder. Her approach integrates computer vision, machine learning, and behavioral psychology to create systems that can objectively measure human behaviors that are often subjective in clinical settings. Analysis of her recent publications reveals a consistent trajectory toward more sophisticated multimodal approaches to healthcare challenges, particularly in pain assessment and mental health diagnostics. Her work increasingly emphasizes interpretable AI models that can translate complex behavioral patterns into clinically meaningful insights, with growing attention to applications for vulnerable populations including infants, elderly patients, and those with craniofacial abnormalities or autism spectrum disorder. Women in AI Awards North America 2023 – AI Researcher of the Year Award Outstanding Reviewer Award at FG 2015 Best Paper award at ACII 2015 Outstanding Paper award at ICMI 2012 Dr. Hammal has secured significant research funding, primarily from the U.S. National Institutes of Health, including an R01 grant for developing a Multimodal Behavioral AI platform for pain assessment and management, and additional grants for automatic pain assessment in older adults with dementia. Her leadership extends to mentoring through her involvement in organizing workshops and conferences that train the next generation of researchers in affective computing and health informatics. As an active leader in her field, Dr. Hammal serves as ACM ICMI Steering Board Committee Member, Associate Editor for IEEE Transactions on Affective Computing and IEEE Transactions on Multimedia, and has organized numerous influential workshops including the International Workshop on Automated Assessment of Pain and Face and Gesture Analysis for Health Informatics. She is set to serve as Program Chair for FG 2025, ACII 2025, and ICMI 2026, demonstrating her growing influence in shaping the future direction of research in multimodal interaction and affective computing.
Prof. Justus Thies is Full Professor for 3D Graphics & Vision at the Technical University of Darmstadt and leads the Neural Capture & Synthesis research group at the Max Planck Institute for Intelligent Systems. His research develops AI methods to capture and synthesize the real world using commodity hardware, focusing on markerless motion capture of faces and bodies, and photorealistic neural rendering. His work has been recognized with the German Pattern Recognition Award, Eurographics Young Researcher Award, and an ERC Starting Grant (all 2024). Recent publications focus on Gaussian-based avatars, neural human reconstruction, and diffusion models for scene synthesis.
Ron Fedkiw is the Canon Professor of Computer Science at Stanford University's School of Engineering. He holds a PhD in Applied Mathematics from UCLA. His research focuses on computational algorithms for applications in computational fluid dynamics, computer graphics, biomechanics, and machine learning. Fedkiw has pioneered techniques for simulating natural phenomena in film and video games, earning two Academy Awards for his contributions to visual effects. He leads the PhysBAM lab and collaborates with industry through consulting roles at Epic Games and former work with Industrial Light & Magic. Education: PhD in Applied Mathematics, UCLA (1996). Notable awards include the National Academy of Science Award, Packard Fellowship, and multiple teaching honors. His lab has graduated 40 PhD students, many of whom have made significant impacts in academia and industry. Research interests span fluid dynamics, cloth simulation, facial animation, and integrating machine learning with physical models. Key contributions include algorithms for two-way fluid-solid coupling, muscle-based facial modeling, and neural network approaches for cloth and deformable bodies. Current projects explore physics-informed machine learning and real-time interactive simulations. Scientific Awards include two Oscars, PECASE, and Okawa Foundation grants. His work bridges computational physics and visual effects, with over 140 research papers and a textbook on level set methods. Advising and grants: Supervised 40 PhD students, securing funding through NSF, ONR, and industrial partnerships. Lab collaborations include SAIL (Stanford AI Lab) and Epic Games. Future work focuses on AI-driven physical simulations and biomedical applications.
Michael J. Black is a Professor and Director at the Max Planck Institute for Intelligent Systems in Tübingen, Germany, where he leads the Perceiving Systems department and serves as Managing Director . He is also an Honorarprofessor at the University of Tübingen 's Faculty of Science . His career spans roles at Brown University (2000-2010), Xerox PARC, and academic-industry collaborations with Amazon and Meshcapade.
Manuel Kaufmann is a Lecturer in the Department of Computer Science at ETH Zürich. His work focuses on advanced 3D human motion capture, sensor-based systems, and computer vision applications. He is affiliated with the Institute of Informatics (inf.ethz.ch) and contributes to research in real-time motion tracking, dataset development, and machine learning integration for human-robot interaction. Research interests include holistic human-scene reconstruction from monocular videos, gaze estimation using EEG signals, and expressive avatar creation. His projects emphasize practical applications in robotics, sports analytics, and biomedical engineering, often leveraging electromagnetic and inertial sensors for high-precision data acquisition. His publications reflect a trend toward multi-modal data fusion, real-world dataset creation (e.g., WorldPose, ARCTIC), and addressing challenges in loose garment modeling (Reloo). These efforts aim to improve markerless motion capture, crowd analysis, and human-robot collaboration. No scientific awards or grants are explicitly listed. He has no documented advisees, though his research may involve collaborations with students or teams. His office is located at OAT X 23, Andreasstrasse 5, Zürich, Switzerland, and contact details include a phone number and professional email.
Andrea Cavallaro is a Full Professor at the École Polytechnique Fédérale de Lausanne (EPFL) and Director of the Idiap Research Institute. He holds dual appointments in the School of Engineering (STI) within the Institute of Electrical Engineering and Measurements (IEM) and the School of Engineering's Education Unit (SEL-ENS). His research focuses on machine learning for multimodal perception, privacy-preserving AI, and autonomous systems. Cavallaro earned his PhD in Electrical Engineering from EPFL in 2002 and has held leadership roles including Director of Research at Queen Mary University of London and Turing Fellow at The Alan Turing Institute. Education: PhD in Electrical Engineering (EPFL, 2002) Leadership: Idiap Director, Affiliate at ELLIS Society Editorial Roles: Editor-in-Chief of Signal Processing: Image Communication (2020–2023), Senior Area Editor for IEEE Transactions on Image Processing Research Interests: Machine learning for audio-visual sensing, privacy in AI, autonomous systems perception, and ethical AI frameworks. Key projects include AlignAI (trustworthy AI alignment) and CORSMAL (multimodal object manipulation). Recent articles explore privacy-aware AI models, adversarial attacks, and multimodal perception systems. His work bridges theoretical advancements with practical applications in robotics, healthcare, and education. Awards include the Royal Academy of Engineering Teaching Prize and IAPR Fellowship. Teaching: Leads courses on deep learning ethics and multimodal AI at EPFL. Advising: Supervises 11 PhD students in areas like privacy-preserving algorithms and autonomous systems. Labs/Teams: Coordinates Idiap’s Audiovisual Intelligence and Learning Lab (LIDIAP) and collaborates on projects like GraphNEx (explainable AI via graph neural networks).
Kevin C. Zhou is an Assistant Professor in the Department of Biomedical Engineering at the University of Michigan. His research focuses on developing high-performance computational optical imaging systems with unprecedented spatiotemporal throughput, integrating advanced optical instrumentation with machine learning-driven algorithms to analyze big data in biology and medicine. His lab specializes in creating imaging systems capable of capturing high-resolution, high-speed, and high-dimensional datasets. Dr. Zhou holds a Ph.D. in Biomedical Engineering from Duke University (NSF GRFP Fellow) and a B.S. in Biomedical Engineering from Yale University (Barry Goldwater Scholar). Prior to joining U-M, he was a Schmidt Science Fellow and postdoctoral researcher at UC Berkeley. Key research areas include: High-throughput microscopy (gigapixel-scale systems) 3D tomographic imaging Light field and Fourier-based imaging modalities Machine learning for image reconstruction and analysis Biomedical applications in cellular/molecular imaging His recent work has advanced technologies like multi-camera array microscopes (MCAM/MCAS) and Fourier light field mesoscopes, achieving video-rate 3D imaging of freely moving organisms. These innovations enable applications in digital cytopathology, behavioral tracking, and high-content biological studies. Notable awards include the NSF Graduate Research Fellowship and Barry Goldwater Scholarship. His research has been featured in top journals and conferences with a focus on advancing optical imaging hardware and computational pipelines.
Marc Erich Latoschik is a Professor in the Department of Human-Computer Interaction at the University of Würzburg, Germany. He previously held roles at Bayreuth University, Bielefeld University, and FHTW Berlin. His research focuses on virtual and augmented reality, embodiment, human-computer interaction, and applications in health, education, and social systems. Education: Completed his PhD in 2001 at Bielefeld University with a thesis on multimodal interaction in virtual reality. Research Interests: Embodied interaction, virtual embodiment, presence and plausibility in VR/AR, avatar design, social virtual reality, health applications (e.g., VR therapy for body image issues), and XR security/privacy. Active in developing frameworks like Reality Stack I/O and MAIL for VR/AR research. Key Projects: ViTraS study on body image exercises in VR, avatars for mass use via smartphone reconstruction, and motion-based biometrics in XR. Collaborates with medical teams on cybersickness detection and emergency training simulations. Labs/Teams: Leads research groups on immersive technologies and social VR applications. Involved in interdisciplinary projects combining HCI, AI, and healthcare.