Khuzaima Daudjee is a Professor and David R. Cheriton Faculty Fellow in the Cheriton School of Computer Science at the University of Waterloo. His research focuses on systems-oriented problems at the intersection of systems and data management, particularly building large-scale systems, storage infrastructure in the cloud, and modern hardware applications. He leads projects in distributed database systems, elastic scaling, and resource optimization. His recent work includes Caerus (geo-replicated transactions), Tiresias (predictive storage), and MorphoSys (automatic physical design metamorphosis). Daudjee has chaired major conferences including ICDE 2026 and serves on editorial boards for VLDB, SIGMOD, and IEEE TKDE journals. His awards include ACM Distinguished Scientist and multiple best paper awards. Educational initiatives include developing distributed systems teaching materials and supervising graduate students across database and distributed systems domains. His industry collaborations involve cloud infrastructure optimization and scalable data processing frameworks.
Erik B. Sudderth is a Professor of Computer Science and Statistics and Chancellor's Fellow at the University of California, Irvine (UCI). He leads the Learning, Inference, & Vision Group and directs multiple research centers, including the UCI Center for Machine Learning and Intelligent Systems and the HPI Research Center in Machine Learning and Data Science. He previously served as an Associate Professor at Brown University. Education: B.S. (summa cum laude) in Electrical Engineering from UC San Diego (1999), M.S. and Ph.D. in EECS from MIT (2002, 2006). His research focuses on statistical methods for scalable machine learning, Bayesian nonparametrics, probabilistic graphical models, and applications in computer vision, AI, and environmental science. Key areas include nonparametric clustering, deep generative models, and particle-based inference algorithms. Research interests span diverse topics: advancing Bayesian nonparametric models for medical time series, scalable variational inference, and AI ethics. Notable contributions include the NET-VISA seismic monitoring system (ISBA Mitchell Prize, 2014), the BNPy toolbox (NSF CAREER Award), and work on diverse particle max-product algorithms for continuous inference. Scientific awards include the NSF CAREER Award, ISBA Mitchell Prize, and recognition as one of "AI's 10 to Watch" (IEEE). He has served as editor for top journals (JMLR, IEEE PAMI) and conference chairs (NeurIPS, CVPR). His work bridges theory and practice, with applications in robotics, climate science, and healthcare. Labs/Teams: UCI Learning, Inference, & Vision Group; UCI Center for Machine Learning; CREATE Technology Center. Grants include NSF funding for visually impaired collaboration tools and soil biogeochemical modeling.
Aditya Parameswaran is an Associate Professor in the Electrical Engineering and Computer Sciences (EECS) department at the University of California, Berkeley. He co-directs the EPIC Data Lab and the Police Records Access project, focusing on simplifying data science at scale through human-in-the-loop systems, LLM-powered tools, and scalable data systems. His research spans database systems, human-computer interaction, and machine learning, with notable contributions in tools like Lux, Modin, and DataSpread. Education : PhD in Computer Science from Stanford University (2013) BTech in Computer Science and Engineering from IIT Bombay (2007) Research Interests : Parameswaran's work centers on empowering end-users with intuitive data tools. Recent projects include LLM-powered systems for document processing (DocETL, TWIX), proactive data systems, and benchmarking frameworks. He emphasizes democratizing data science through low/no-code solutions and improving production ML workflows. Articles Trends : His recent work (2023–2025) prioritizes LLM integration into data systems, focusing on robust pipelines, assertion generation (SPADE), and debugging tools (RAGGY). Earlier contributions include visualization recommendation (Lux), scalable dataframes (Modin), and spreadsheet optimization (DataSpread). Awards : Recipient of the VLDB Early Career Award (2019), Sloan Research Fellowship (2020), NSF CAREER Award (2017), and multiple best paper/demonstration awards at top venues like SIGMOD and VLDB. Advising & Grants : Guides over 20 PhD/postdoc alumni, many now in academia (e.g., Madelon Hulsebos at CWI) and industry leadership roles. Active in securing grants (e.g., NSF, Army Research Office) and industry partnerships (e.g., Snowflake, LangChain). Labs/Teams : Leads the EPIC Data Lab, focusing on agentic data systems, and co-founded Ponder (acquired by Snowflake). Collaborates on the Police Records Access initiative, building transparency tools for public records.
Professor Hanumant Singh leads the Electrical and Computer Engineering department at Northeastern University, with a joint appointment in Mechanical and Industrial Engineering , and serves as Program Director for the Master of Science in Robotics. He earned his Ph.D. from MIT/WHOI Joint Program in 1995 and has conducted over 60 expeditions globally, focusing on marine geology, polar studies, and coral reef ecology. His research emphasizes field robotics , including SLAM, underwater manipulation, and imaging in extreme environments. He developed the Seabed AUV and Jetyak ASV , widely used in scientific research. His labs include the Field Robotics Lab and the Institute for Experiential Robotics . Research Interests: Machine Learning for Fisheries SLAM with dynamic objects Underwater imaging and manipulation Autonomous surface and aerial systems Polar and marine robotics Awards: ICRA Best Student Paper Award, IEEE Oceanic Engineering Society Distinguished Faculty Award (2025), Lifetime Achievement Award (2022), and IEEE Fellow status. His work has been featured in Nature Geoscience , Polar Biology , and media outlets like WGBH. Students & Collaborations: Advises students like Srinidhi Pattala (MS Robotics) and Dennis Giaya (PhD Computer Engineering). Collaborates with institutions on projects such as Antarctic sea ice thickness estimation and deep-sea submersible missions.
Andy Pavlo is an Associate Professor with Indefinite Tenure in the Computer Science Department at Carnegie Mellon University's School of Computer Science. He is an active member of the CMU Database Group and the Parallel Data Laboratory, where he leads research in database management systems with a focus on self-driving architectures, transaction processing, and large-scale analytics. His work bridges academic research and industry applications through projects like NoisePage, OtterTune (which he co-founded and served as CEO before it ceased operations), and Peloton. Dr. Pavlo's research interests span database management systems with particular emphasis on autonomous database architectures that can self-tune and optimize without human intervention. His work explores transaction processing systems that can handle high-throughput workloads while maintaining consistency, and large-scale data analytics techniques that efficiently process massive datasets. He has made significant contributions to query optimization, database extensibility, and automatic database tuning using machine learning techniques. His recent work on database extensibility revealed critical issues in PostgreSQL's extension ecosystem, showing that approximately 16% of extensions are incompatible with at least one other extension due to API violations and memory errors. His research output demonstrates a consistent focus on practical database systems challenges, with recent publications examining database extensibility, user-defined function optimization, and the cyclical nature of database research. The articles show a strong trend toward making database systems more autonomous, with increasing integration of machine learning techniques for automatic tuning and optimization. His work often combines deep theoretical analysis with practical implementation in open-source systems. Dijkstra Award 2024 for contributions to database systems research Dr. Pavlo actively mentors graduate students, with current advisees including Wan Shen Lim, William Zhang, and Sam Arch (co-advised with Todd Mowry). His former students have gone on to successful careers in both industry and academia. He has secured significant research funding through CMU's affiliate program with major database companies including ClickHouse, DataStax, dbt, Firebolt, MotherDuck, RelationalAI, SingleStore, Spiral, PingCAP/TiDB, Yellowbrick, and Yugabyte. His research is supported by these industry partnerships and likely includes NSF funding given his active participation in the database research community. At CMU, Dr. Pavlo leads the Database Group and organizes several seminar series including "SQL or Death," "Database Building Blocks," and "ML⇄DB Technical Talks." These seminars bring together researchers and practitioners to discuss cutting-edge developments in database systems. He also runs a summer research internship program that has attracted students for multiple consecutive years, indicating a strong research group with ongoing projects and funding.
Steve Mussmann serves as an Assistant Professor in the School of Computer Science at the Georgia Institute of Technology, where he joined in Fall 2024. His research centers on data-centric machine learning, with emphasis on active labeling, data selection, and adaptive experimental design methodologies. He maintains active collaborations through Georgia Tech's Foundations of AI (FoAI) and ML@GT research groups. Mussmann earned his PhD in Computer Science from Stanford University in 2021 under Percy Liang's supervision, following a BS in Math, Statistics, and Computer Science from Purdue University in 2015. His professional trajectory includes a machine learning researcher role at Coactive AI and an IFDS postdoctoral fellowship at the University of Washington's Paul Allen School of Computer Science and Engineering. His research program investigates theoretical and practical aspects of data efficiency in machine learning systems, particularly focusing on active learning frameworks, statistical properties of data algorithms under concept drift, and task specification via prompts or demonstrations. Current projects address challenges in label-efficient training of large language models and multimodal dataset development. Analysis of his 15 most recent publications reveals a consistent focus on advancing data-centric methodologies, with increasing emphasis on large-scale applications like multimodal datasets and language model fine-tuning. His work bridges theoretical guarantees in experimental design with practical frameworks like LabelBench for benchmarking label efficiency. Mussmann has received recognition through the IFDS postdoctoral fellowship. His contributions to the field include foundational work on active learning theory and data selection algorithms. IFDS postdoctoral fellow He currently advises five graduate students including PhD candidates Kangping Hu (CS) and Hangyu Zhou (ML), alongside MS students Kabir Kang and Kalp Vyas, and undergraduate Saloni Bedi. Former advisee Wei-Liang (Edison) Liao completed BS research under his supervision. His teaching portfolio includes graduate courses CS 7545 (Machine Learning Theory) and CS 8803-DML (Data-centric Machine Learning). Mussmann operates within Georgia Tech's Foundations of AI initiative and ML@GT collective, which provide infrastructure for large-scale data-centric research. His lab develops open-source tools like LabelBench for reproducible evaluation of data selection techniques, with ongoing projects exploring video data exploration systems and adaptive finetuning frameworks for foundation models.
Fabian Suchanek is a full professor at Institut Polytechnique de Paris, specifically affiliated with Télécom Paris. He leads research in the Data, Intelligence, and Graphs (DIG) team within the Computer Science department. His academic career focuses on bridging artificial intelligence with structured knowledge representations. Suchanek's research interests span artificial intelligence, knowledge bases, and natural language processing, with particular emphasis on knowledge graph construction , rule mining , knowledge-based language models , and explainable AI . His work demonstrates how structured knowledge can enhance machine learning systems, particularly large language models, by providing factual grounding and interpretability. The research group he leads develops practical systems that address real-world knowledge management challenges. His recent publications showcase a strong trajectory in knowledge-intensive AI, with notable contributions to knowledge graph completion, rule mining techniques, and neural approaches to knowledge base validation. The research demonstrates increasing integration between symbolic and neural approaches to AI. Best Student Paper Award at KR 2024 for work on contextual reasoning Best Demo Award of IJCAI 2024 for rule mining in knowledge graphs French Open Research Award for the YAGO project Best Paper Award of ESWC 2021 for Neural Knowledge Base Repairs Suchanek has secured significant research funding, evidenced by his active recruitment of PhD students for knowledge-based language model research. He has held visiting positions, including at Nanyang Technological University (June-September 2023), and is recognized internationally through keynote invitations such as the Singapore ACM SIGKDD Symposium 2023. He has deliberately stepped back from administrative duties at Institut Polytechnique de Paris to focus on research. His laboratory maintains strong industry connections through open-source software projects including the YAGO knowledge base, AMIE for rule mining, STACI for explainable AI, and several other tools that have become standard in knowledge representation research.
Enamul Hoque Prince is an Associate Professor and Director of the School of Information Technology at York University. He leads the Intelligent Visualization Lab, funded by the Canada Foundation for Innovation (CFI) and Ontario Research Funds (ORF). He holds a PhD in Computer Science from the University of British Columbia and completed postdoctoral work at Stanford University. His research integrates information visualization, human-computer interaction (HCI), and natural language processing (NLP) to address information overload challenges. Dr. Prince's educational background includes a PhD from UBC, an MSc from Memorial University of Newfoundland, and a BSc from Chittagong University of Engineering & Technology. He has conducted research at institutions like Tableau Software and the Qatar Computing Research Institute and serves on committees for top conferences like ACL and IEEE Vis. His work is supported by grants from NSERC, CFI, and others. Research interests focus on NLP-driven visual analytics, user-adaptive visualization, and accessible interfaces. Notable projects include Evizeon (natural language interfaces for visual analytics), ConVisIT (topic modeling for online conversations), and CIDER (concept-based image search). Publications highlight trends in multimodal systems, chart comprehension, and accessibility. Key awards include the NSERC Discovery Grant (2019) and a Best Paper Honorable Mention at DIS 2021. He supervises graduate and undergraduate students in areas like visualization, NLP, and HCI. Labs and collaborations emphasize interdisciplinary approaches, with the Intelligent Visualization Lab advancing tools for data exploration and user-centered design. Teaching includes courses on design principles and information visualization.
Georgia Gkioxari is an Assistant Professor in the Division of Computing and Mathematical Sciences at Caltech , with a part-time affiliation at Meta AI . Her work focuses on extending visual perception models through advanced 2D and 3D representation learning, spatial reasoning, and generative models. Education: Not explicitly mentioned in the text Research interests span 3D perception , spatial reasoning , and vision-language integration , with projects like Visual Agentic AI for Spatial Reasoning and Token-by-Token Multimodal Alignment . Her publications emphasize 3D object detection , reconstruction , and generative modeling techniques including diffusion models and transformers . Scientific recognition includes the Meta LLM Evaluation Research Grant , Okawa Research Grant , Google Faculty Scholar Award 2024 , and Amazon Research Award . She teaches courses like Large Language & Vision Models (EE/CS 148) and Learning & 3D (CS 101) at Caltech. Labs & Teams: Leads Glab with members including Ilona Demler, Ziqi Ma, and Damiano Marsili
Zachary Ives is the Adani President's Distinguished Professor and Department Chair of the Computer and Information Science Department at the University of Pennsylvania. He holds affiliations with the ASSET Center for Safe, Explainable and Trustworthy AI, the Warren Center for Network and Data Science, the Center for Neuroengineering and Therapeutics, and serves as a Distinguished Research Fellow at the Annenberg Center for Public Policy. His research focuses on data integration and sharing, data provenance and trustworthiness, and machine learning systems. He develops data science platforms at the intersection of databases, machine learning, and distributed systems, with applications in Web question answering and scientific domains like genetics and neuroscience. His work addresses fundamental challenges in integrating heterogeneous data, ensuring trustworthy results, and facilitating collaborative data science. His recent publications demonstrate a strong focus on data lakes, learned database systems, fine-grained provenance, and question answering systems. These works span top conferences including SIGMOD (where his paper was selected as Best Paper in 2024), VLDB, ACL, and PODS, showing the breadth of his contributions across database systems, natural language processing, and data management. NSF CAREER award recipient Fellow of the ACM Christian R. and Mary F. Lindback Foundation Award for Distinguished Teaching IEEE Technical Committee on Data Engineering Education Award SIGMOD Best Paper Award ICDE 2013 ten-year Most Influential Paper award As Department Chair, Ives has overseen significant departmental growth, hiring 25 new faculty since 2018. He advises numerous PhD students and postdocs, and maintains extensive collaborations across Penn and with external institutions. His research has been funded by NSF, NIH, DARPA, Google, Amazon, and other organizations. He has developed courses including NETS 212 'Scalable and Cloud Computing' and teaches Big Data Analytics. His research group, the Penn Database Group, works on projects including data lake management, data provenance, and collaborative data science platforms. His work with neuroscientists on seizure prediction has received significant attention, including a competition with 504 teams achieving 82% accuracy.
Joy Arulraj is an Associate Professor in the School of Computer Science within the College of Computing at Georgia Institute of Technology. His research focuses on data systems, machine learning, and database systems, with a particular emphasis on video analytics and adaptive query processing. He leads the Data Systems and Analytics Group and is developing the EVA AI-Relational Data System. Dr. Arulraj's research interests span data systems, machine learning, database systems, video analytics, and adaptive query processing. His work centers on developing systems that efficiently process complex queries, particularly for video analytics and machine learning workloads. He has made significant contributions to GPU database systems, non-volatile memory database management, and adaptive query processing techniques. His research often bridges the gap between theoretical database principles and practical implementations for modern hardware architectures. His recent publications show a strong trend toward video analytics systems, adaptive query processing for machine learning workloads, and GPU-accelerated database systems. The EVA system represents a major focus of his recent work, providing end-to-end exploratory video analytics capabilities. His research also addresses fundamental database concepts like buffer management, query optimization, and storage management, adapting these principles for modern hardware and application requirements. Dr. Arulraj has advised numerous graduate students including Pramod Chunduri, Gaurav Tarkok Kakkar, Jiashen Cao, and Sayan Sinha. His graduated students have gone on to work at companies like ServiceNow, Meta Research, and the Korean Army. He actively teaches database system courses at Georgia Tech, including Database System Implementation (CS 4420/6422) and Advanced Database System Implementation (CS 4423/6423), where students build database systems from scratch using C++ and the BuzzDB framework. He maintains an active research program with consistent publication output across top database and systems conferences. His work spans from theoretical database principles to practical system implementations, with a recent emphasis on video analytics, machine learning integration with database systems, and leveraging modern hardware like GPUs and non-volatile memory for database applications.
Hamed Zamani is an Associate Professor at the Manning College of Information and Computer Sciences (CICS) at the University of Massachusetts Amherst, where he also serves as Associate Director of the Center for Intelligent Information Retrieval (CIIR). He joined UMass Amherst in 2020 after working as a researcher at Microsoft. His research focuses on designing and evaluating statistical and machine learning models for information access systems, including search engines, recommender systems, and question answering. Education: PhD in Computer Science, University of Massachusetts Amherst MS in Computer Engineering, University of Tehran BS in Computer Engineering, University of Tehran Zamani's current research explores neural information retrieval, conversational search, and retrieval-enhanced machine learning. He develops efficient neural models for core IR tasks and emerging areas like conversational information seeking. His work bridges information retrieval with large language models to enhance capabilities in understanding complex queries and generating relevant responses. His recent publications demonstrate a strong focus on retrieval-augmented generation, personalized information access, and efficient neural ranking models. There's a clear trend toward integrating large language models with information retrieval systems, optimizing multi-agent frameworks, and developing evaluation metrics for generative AI applications in search contexts. Scientific Awards: NSF CAREER Award ACM SIGIR Early Career Excellence in Research & Community Engagement Awards (2023) UMass CICS Outstanding Dissertation Award Paper awards at SIGIR (2022, 2023, 2024), CIKM (2020), ICTIR (2019) Microsoft Research Award (AI and New Future of Work program) Amazon Research Award (Optimization of Retrieval-Enhanced ML Models) Zamani actively advises PhD students and postdoctoral researchers, with his students receiving prestigious awards including NSF Graduate Research Fellowships and SIGIR Best Paper awards. He leads the CIIR Talk Series, hosting IR researchers to share recent findings. His Alexa Prize TaskBot Challenge team was selected for two consecutive years, advancing task-oriented dialogue systems. He directs research at the Center for Intelligent Information Retrieval (CIIR), where he oversees projects in neural retrieval models, conversational AI, and retrieval-augmented generation. The center serves as a hub for developing next-generation information access systems with industry and academic collaborators.
José F. Martínez holds the Lee Teng-hui Professorship of Engineering at Cornell University's College of Engineering, where he leads research in computer architecture and systems. His roles include Vice Chair of ACM SIGARCH, IEEE Fellow, and past leadership roles in ISCA, MICRO, and IEEE Computer Society journals. He has received prestigious awards such as the NSF CAREER Award, IBM/Qualcomm Faculty Awards, and multiple teaching accolades including Tau Beta Pi Professor of the Year (2011). Education: Licenciado en Informática de Sistemas (1996), Universidad Politécnica de Valencia M.S. (1999) and Ph.D. (2002) in Computer Science, University of Illinois at Urbana-Champaign His research focuses on computer architecture , including microprocessors, multiprocessors, memory subsystems, embedded systems, and processing-in-memory (PIM) innovations. He co-leads the DARPA/SRC ACE Center for Evolvable Computing and NSF's CROPPS Center, and previously co-founded Cornell's Institute for Digital Agriculture (CIDA). His work bridges hardware-software co-design and sustainable computing. Key Contributions: Advancing energy-efficient multiprocessor architectures Developing PIM frameworks like PUMICE and Membrane Market-based resource allocation systems (e.g., XChange, ReBudget) His advising spans over 15 graduate students and has guided Merrill Presidential Scholars. Current research explores evolvable computing paradigms and AI hardware acceleration, reflecting his dual focus on foundational architecture and applied systems innovation.
James Glass is a Senior Research Scientist at the Massachusetts Institute of Technology (MIT) and heads the Spoken Language Systems Group within MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). He is also affiliated with the Harvard-MIT Division of Health Sciences and Technology. His research spans automatic speech recognition, multimodal learning, and spoken language understanding, with applications in healthcare and video analysis. Education: SM and PhD in Electrical Engineering and Computer Science from MIT His work focuses on paralinguistic speech analysis, health markers in speech, and the intersection of speech and natural language processing. Recent trends emphasize audio-visual alignment, recursive reasoning, and AI applications in cognitive disorder diagnosis. Scientific awards include IEEE Fellow, ISCA Fellow, and Associate Editor for IEEE Transactions on Pattern Analysis and Machine Intelligence. His group explores unsupervised learning, speaker verification, and social text analysis. James leads the Spoken Language Systems Group at CSAIL, collaborating with institutions like IBM and Harvard-MIT Division of Health Sciences and Technology. His research integrates vision-language models, neural audio codecs, and self-supervised frameworks.
Madelon Hulsebos is a Researcher at CWI in Amsterdam, where she leads the Table Representation Learning (TRL) Lab and contributes to the Database Architectures group. She is also a faculty member of the European Laboratory for Learning and Intelligent Systems (ELLIS) Amsterdam unit. Her career bridges academia and industry, including a postdoctoral fellowship at UC Berkeley and prior industry experience in automating data analysis pipelines with ML. Education : PhD in Computer Science (University of Amsterdam, 2023), with research at Sigma Computing and MIT; Postdoctoral Fellow (UC Berkeley, 2024). Her research focuses on establishing tabular data as a key AI modality through Table Representation Learning , generative models for relational data, and robust systems for data analysis. Key interests include: Relational Table Embeddings LLMs for QA/text2SQL and data wrangling Retrieval over Data Lakes and Databases Agentic Systems for Data Science Democratizing insights from structured data Recent work highlights trends in benchmarking table retrieval (TARGET), semantic column detection (AdaTyper, Sherlock), and large-scale tabular data curation (GitTables, SchemaPile). These projects address challenges in metadata utilization, data lake search, and end-to-end systems for structured data. She has secured significant funding, including the NWO AiNed Fellowship Grant ($1M) for her 5-year DataLibra project. Madelon organizes workshops at NeurIPS , SIGMOD , and ACL , and reviews for top venues like VLDB and NeurIPS. Scientific Awards : NWO AiNed Fellowship Grant ($1M) She actively mentors students and collaborates on European AI initiatives, including monthly TRL seminars and workshops. Her lab's tools (GitTables, TARGET) are widely adopted for training foundation models on tabular data.