Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 18, 2026 13:15 - 13:40 CEST
Frontier labs are spending billions making agents better at SWE-Bench. But how much of your engineering work actually looks like SWE-Bench? At Datadog we kept seeing agents that crushed public benchmarks fail on our codebase: missing our conventions, reaching for the wrong internal libraries, technically correct but doing the work the wrong way.

To get past demos and gut feel, we built an evaluation platform that measures agents on tasks drawn from our real work, and gave our platform teams a way to encode best practices as evals. Teams shipping skills, steering docs, agent harnesses, and MCP servers can now see whether their changes actually moved the needle.

In this talk I'll share how SOTA and open-weight models actually compare on real work, what their cost-performance profiles look like, tooling decisions that can shift token usage by 10% or more, and how a surprisingly small eval suite can produce stable signal.

You'll leave with a clearer way to think about model choice as a tradeoff between performance you actually need and tokens you're willing to spend, and a sharper sense of what makes an eval keep paying off over time instead of becoming a one-off exercise.
Speakers
avatar for Ville Hellman

Ville Hellman

Staff Engineer, Datadog
Ville is a Staff Engineer in Datadog's AI DevX group, where he builds evaluation infrastructure for the agents and tooling Datadog engineers use every day. He writes on AI-augmented engineering, AI literacy, and developer experience to make what's coming next easier for others to... Read More →
Friday September 18, 2026 13:15 - 13:40 CEST
G104 + G105 (Level 1)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link