Parth Suresh
Member of Technical Staff · Datology AI
I’m a Member of Technical Staff at Datology AI, where I work on synthetic data generation for web-scale and long-context data.
Previously, I was a Software Engineer (Machine Learning) at Meta Reality Labs, working on synthetic data, data curation, and benchmarks for language and multimodal models, including in egocentric and wearable settings. Before that I was an ML Research Engineer at Scale AI, focused on reasoning, synthetic data generation, and LLM judges. Earlier at Meta I worked at the intersection of developer productivity, software engineering, and machine learning.
The question I keep coming back to is how we turn messy, real-world data into reliable signals for training and evaluating large models.
News
| Aug 10, 2026 | CRAG-MM won the Best Paper Award at KDD 2026 in the Datasets and Benchmarks track! |
|---|---|
| Jul 20, 2026 | Joined Datology AI as a Member of Technical Staff, working on synthetic data generation for web and long-context data. |
| Jun 03, 2026 | Our paper Plan, Watch, Recover, a benchmark and architectures for proactive procedural assistance, is now on arXiv. |
| Oct 30, 2025 | Released CRAG-MM, a multi-modal multi-turn RAG benchmark for wearable / egocentric settings - and the foundation for KDD Cup 2025. |
| Dec 10, 2024 | Our paper on balancing cost and effectiveness of synthetic data generation strategies for LLMs was accepted to the FITML Workshop at NeurIPS 2024. |
Selected publications
-
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG BenchmarkIn Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2026