Agent evaluation new paradigm: from single benchmark to multi-dimensional safety alignment testing
Anthropic and OpenAI have released new agent evaluation suites, expanding from simple code completion accuracy to multi-turn interaction safety, tool call hallucination, permission abuse detection, and more. Agent eval enters the safety alignment era.