Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 18, 2026 15:45 - 16:10 CEST
LLM agents are reaching production faster than teams can evaluate them. A data-analysis agent that runs the right query but reports the wrong number, or returns the right number via a trajectory full of fabricated tool calls, passes superficial testing and fails in production.

This talk walks through evaluating such an agent end-to-end. Our running example: a data-analysis agent answering questions over a business dataset. We show how to grade three dimensions that agent evaluation requires and single-shot LLM evaluation ignores: final response, trajectory, and state changes.

We cover the full lifecycle:

1. Bootstrapping evaluation from 50 hand-reviewed examples when you have no labels.
2. Aligning an LLM-as-a-judge to human judgment with the same rigor you'd apply to outsourced annotators: dev/test splits, inter-rater agreement, Cohen's kappa.
3. Scaling to continuous online evaluation with CI integration, error analysis, and prompt optimization driven by natural-language feedback.

We also cover what we got wrong in earlier iterations and what we'd do differently today.
Speakers
avatar for Bauke Brenninkmeijer

Bauke Brenninkmeijer

AI Research Engineer, Orq.ai
Bauke is an AI Research Engineer with a background in data science and computer science. After working at several startups, I spent 5 years building ML systems at ABN AMRO and ING — from real-time streaming frameworks to RAG-based document processing.Now at orq.ai, I focus on AI... Read More →
Friday September 18, 2026 15:45 - 16:10 CEST
G106 + G107 (1st Floor)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link