Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 18, 2026 12:05 - 12:30 CEST
A three-layer evaluation strategy for AI agents and their tools, from experience operating across multiple LLM providers and runtimes.

Offline evaluation: designing benchmark suites with curated requests, expected tool selections, and arguments. Computing precision, recall, F1 scores, confusion matrices for tool mix-ups, and argument hallucination rates to pinpoint description problems.

End-to-end benchmarks: multi-tool flows where the agent chains several calls to complete a task, catching integration regressions that single-tool evaluation misses.

Production A/B experiments: a case study of tool search experiments across OpenAI and Anthropic through staged rollouts. Challenges we hit: caching bugs under real traffic, tool discovery failures, and data skew making early results inconclusive. How we decided whether to advance, pause, or roll back.

These layers compensate for each other's blind spots, forming a testing pyramid that enabled us to safely ship changes to MCP and agent across model providers. Attendees leave with a reusable playbook for testing MCP servers, agents, and running A/B experiments.
Speakers
avatar for Ksenia Bobrova

Ksenia Bobrova

Senior Software Engineer, GitHub
I'm currently focusing on building AI agents and tooling around it.
Friday September 18, 2026 12:05 - 12:30 CEST
G104 + G105 (Level 1)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link