CV
Contact Information
| Name | Jiarui (Jerry) Yan |
| Professional Title | MS in Computational Data Science, Carnegie Mellon University |
| jerryy2@cs.cmu.edu |
Professional Summary
Master’s student at Carnegie Mellon working on LLM pre-training data, post-training, reinforcement learning, and agentic systems. Previously a software engineer intern at Google.
Education
-
2025 - 2027 Pittsburgh, PA
MS
Carnegie Mellon University, School of Computer Science
Computational Data Science
- Large Language Models, Advanced NLP, Multimodal Machine Learning, Deep Learning, Cloud Computing, Distributed Systems
-
2021 - 2025 Seattle, WA
BS
University of Washington, Information School
Informatics: Data Science (major), Applied Mathematics (minor)
- Magna Cum Laude (top 3.5%); Dean’s List every term.
- Coursework: Machine Learning, NLP, Numerical Analysis, Probability I & II, Linear Algebra, Methods in Data Science.
Research
- LLM Pre-training & Post-training (01/2026 – present) — Research Assistant with Prof. Sanmi Koyejo (Stanford) and Graduate Researcher with Prof. Chenyan Xiong (CMU); manuscripts in preparation. Developing scaling laws that predict how much RL post-training will improve a model from measurable properties of the training environment. Built a model-based HTML-to-text refiner (Qwen3-1.7B, SFT + GRPO on vLLM/Verl) lifting token retention from 40.5% to 54.7% on DCLM CommonCrawl; models pretrained on the refined tokens beat the best heuristic baseline by 3–4% on DCLM Core.
- Agentic Systems & Agent Planning (12/2025 – 06/2026) — Graduate Researcher with Prof. William W. Cohen (CMU, Google DeepMind) and Prof. Yiming Yang (CMU). Built a modular agentic-LLM framework centered on pseudo-tools and ran the first systematic study across 19 evaluation suites showing static workflows beat dynamic planners (0.64 vs 0.56 average correctness at half the cost). Introduced a dataset of 4,672 paired human/agent Kaggle trajectories exposing a structural planning gap in agents; a prompt-level harness improved leaderboard score on most of 7 competitions (released on HuggingFace, CC-BY-4.0).
- Earlier, University of Washington (2023 – 2024) — transfer-learning matrix completion for multi-source recommendation with Prof. Wanning Chen (RMSE 0.6 → 0.35 on NetEase data); large-scale GitHub / Hugging Face mining of fork behavior in open source with Prof. Jane Tan and Prof. Yong Tan.
Teaching
- Teaching Assistant, University of Washington (2023 – 2024) — INFO 498 Text Mining & Analytics and INFO 330 Databases & Data Modeling with Prof. Lucy Lu Wang; designed NLP assignments and led weekly SQL recitations for 35 students.
Skills
Languages: Python, C++, Java, TypeScript, JavaScript, SQL, R, Shell
ML / LLM: PyTorch, vLLM, Verl, Hugging Face, TensorFlow, Pydantic-AI, LiteLLM, LangChain, CUDA
Data & Infra: Spark, Hadoop, HDFS, Kafka, Airflow, AWS, GCP, Docker, Kubernetes