Publications

Peer-reviewed publications and preprints. Google Scholar →

Peer-Reviewed

AgentSuite: Toward More Reliable Agent Evaluation with a Component-Based Benchmark Auditing Pipeline

H. Suh*, S. Lee*, B. Ji*, R. Khare, B. Khan, H. Kim, Tianyi Zhang, et al. (* equal contribution)

ICML 2026 Paper

A component-based pipeline that audits validity flaws in LLM-agent benchmarks — decomposing tasks into user, environment, ground-truth, and evaluation components. It matches expert judgments at 0.79–0.87 F1 across 6 benchmarks and reshuffles 63% of model rankings on a 30-LLM leaderboard.

Preprints

Preprints and manuscripts under review will appear here.