How benchmarks and evaluation strategies have changed as we move from LLMs → agents, and what lies ahead.
This survey traces the chronological development of AI agent benchmarks from 2018 to 2025, examining how evaluation methodologies have adapted to measure increasingly sophisticated capabilities.
Read the full article on Substack: Evaluation for Agentic AI