Tianyi
Zhang

I work on trustworthy AI, end to end — tracking where data, models, and metrics drift from the truth they stand for, and how those gaps compound.

Emory ’25  |  Georgia Tech ’26
ML Engineer Intern · LinkedIn Feed AI

An abstract node graph: connected points with two highlighted nodes, evoking structure, connection, and inspection.
Evaluation · Robustness · Integrity
Namcha Barwa panorama, Tibet — photo by Tianyi Zhang
slide to reveal the time-lapse ▸
Namcha Barwa 南迦巴瓦峰

About

My through-line is proxy fidelity across the ML pipeline. Training data is a proxy for the real distribution; a model's learned objective is a proxy for what you actually wanted; an evaluation metric is a proxy for whether it's genuinely good and safe. The same gap shows up at all three — and it compounds: a proxy-corrupted dataset can train a proxy-gaming model that a proxy-blind eval scores as clean. I study those failures at the interfaces, across the whole chain, not inside any single stage. Where the gap shows up:

  • Data — is a benchmark a faithful proxy for real task quality? (AgentSuite, ICML 2026)
  • Model — did it learn the objective we meant, or one it can game? (retrieval & recommender systems)
  • Eval — does a metric or safety monitor actually track quality and safety? (AgentSuite's LLM-judge, vs. expert labels)

I'm finishing an MS in Computer Science at Georgia Tech (Dec 2026) and spending the summer as an ML engineer intern at LinkedIn (Feed AI). Across all of it, the same question keeps recurring — whether the numbers we report mean what we claim they mean.

News

  • 2026Joined LinkedIn's feed recommendation team (Feed AI) for the summer as an ML engineer intern.
  • 2026🎉 AgentSuite, a component-based benchmark-auditing pipeline for LLM agents, accepted at ICML 2026.
  • 2025🎉 Oral presentation at IEEE EMBC 2025 — a state-based Transformer for fMRI.
  • 2025🎉 First-authored paper on causal brain-connectivity analysis published at AIME 2025.