Jiarui Yan

MS in Computational Data Science, School of Computer Science, Carnegie Mellon University

prof_pic.jpg

Pittsburgh, PA

jerryy2 [at] cs.cmu.edu

I am a master’s student in Computational Data Science at Carnegie Mellon University, working with Chenyan Xiong on pre-training data, William W. Cohen on agentic systems, and Yiming Yang on agent planning. I am also a research assistant with Sanmi Koyejo at Stanford.

My research asks what actually makes a language model better: which data is worth training on, which environments make RL pay off, and why agents still fail on long-horizon work.

Previously I interned at Google and graduated magna cum laude from the University of Washington. I am looking for research scientist / engineer roles starting 2027.

news

Sep 24, 2026 Released TraceML: the paper, a project page, and an open-source toolkit for reading any ML-agent run against human Kaggle practice.
Aug 15, 2026 Wrapped up my software engineering internship at Google, where I built and shipped an LLM agent for Workspace SRE access authorization.
Jul 10, 2026 TraceML was accepted to the WAB workshop at COLM 2026.
May 18, 2026 Started as a research assistant with Prof. Sanmi Koyejo at Stanford, working on scaling laws for RL environments.
Jan 12, 2026 Started working with Prof. Chenyan Xiong on model-based pre-training data refinement.

selected publications

  1. arXiv
    Learning to Construct Practical Agentic Systems
    Aditya Kumar, Zhihan Lei, Jiarui Yan, and 5 more authors
    Under review, 2026
    Equal contribution: A. Kumar, Z. Lei, J. Yan
  2. NeurIPS
    TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
    Jiarui Yan, Weiwei Sun, Sijie Li, and 2 more authors
    In Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets, 2026
    Equal contribution: J. Yan, W. Sun. Also at the COLM 2026 Workshop on Agent Behavior (WAB)
  3. In prep
    HTML2Text: A Unified Model-Based Refiner for Pre-training Data
    Jiarui Yan, Zichun Yu, Shlok Sanghvi, and 1 more author
    Manuscript in preparation, 2026
    Equal contribution: J. Yan, Z. Yu