Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Type: Evaluation & Testing clear filter
Friday, September 18
 

12:05 CEST

Testing Agents and Their Tools: Offline Evaluation, Synthetic Tasks, and A/B Experiments - Ksenia Bobrova, GitHub
Friday September 18, 2026 12:05 - 12:30 CEST
A three-layer evaluation strategy for AI agents and their tools, from experience operating across multiple LLM providers and runtimes.

Offline evaluation: designing benchmark suites with curated requests, expected tool selections, and arguments. Computing precision, recall, F1 scores, confusion matrices for tool mix-ups, and argument hallucination rates to pinpoint description problems.

End-to-end benchmarks: multi-tool flows where the agent chains several calls to complete a task, catching integration regressions that single-tool evaluation misses.

Production A/B experiments: a case study of tool search experiments across OpenAI and Anthropic through staged rollouts. Challenges we hit: caching bugs under real traffic, tool discovery failures, and data skew making early results inconclusive. How we decided whether to advance, pause, or roll back.

These layers compensate for each other's blind spots, forming a testing pyramid that enabled us to safely ship changes to MCP and agent across model providers. Attendees leave with a reusable playbook for testing MCP servers, agents, and running A/B experiments.
Speakers
avatar for Ksenia Bobrova

Ksenia Bobrova

Senior Software Engineer, GitHub
I'm currently focusing on building AI agents and tooling around it.
Friday September 18, 2026 12:05 - 12:30 CEST
G104 + G105 (Level 1)

12:40 CEST

Beyond Vibe-Testing: Engineering Deterministic Agent Skills - Shuva Jyoti Kar, Palo Alto Networks
Friday September 18, 2026 12:40 - 13:05 CEST
The AI ecosystem suffers from a critical engineering immaturity: deploying stochastic models via manual "vibe checks." Operating autonomous agents at enterprise scale requires abandoning ad-hoc observation for strict, distributed systems rigor. This session introduces a deterministic, CI/CD-native evaluation architecture for Agent Skills, shifting from indeterministic to reliable software execution.

By adhering to the formalized capability standards defined by agentskills.io, we will deconstruct the transition from subjective testing to hermetic, code-driven audits. Attendees will learn to engineer scenario matrices that enforce strict cognitive boundaries via negative testing—guaranteeing agents safely reject out-of-scope triggers. We will demonstrate isolating execution within sandboxed environments to capture pristine telemetry: deterministic tool-call structures, system exit codes, and exact token utilization.

Crucially, we address the anti-pattern of relying on "LLM-as-a-judge" for critical path assertions. Instead, we architect a framework grading system invariants via AST parsing and JSON Schema enforcement to achieve instantaneous, hallucination-immune evaluation.
Speakers
avatar for SHUVA JYOTI KAR

SHUVA JYOTI KAR

Principal Engineer, Palo Alto Networks
Shuva is a Principal Engineer at Palo Alto Networks building secure enterprise AI platforms. An author of two upcoming books: Engineering the Data Agent Control Plane (O'Reilly) and Agent Skills in Action (Manning), an open-source contributor and former OpenDaylight committer, his... Read More →
Friday September 18, 2026 12:40 - 13:05 CEST
G104 + G105 (Level 1)

13:15 CEST

From Vibes To Data: Evaluating Agents on Your Real Work - Ville Hellman, Datadog
Friday September 18, 2026 13:15 - 13:40 CEST
Frontier labs are spending billions making agents better at SWE-Bench. But how much of your engineering work actually looks like SWE-Bench? At Datadog we kept seeing agents that crushed public benchmarks fail on our codebase: missing our conventions, reaching for the wrong internal libraries, technically correct but doing the work the wrong way.

To get past demos and gut feel, we built an evaluation platform that measures agents on tasks drawn from our real work, and gave our platform teams a way to encode best practices as evals. Teams shipping skills, steering docs, agent harnesses, and MCP servers can now see whether their changes actually moved the needle.

In this talk I'll share how SOTA and open-weight models actually compare on real work, what their cost-performance profiles look like, tooling decisions that can shift token usage by 10% or more, and how a surprisingly small eval suite can produce stable signal.

You'll leave with a clearer way to think about model choice as a tradeoff between performance you actually need and tokens you're willing to spend, and a sharper sense of what makes an eval keep paying off over time instead of becoming a one-off exercise.
Speakers
avatar for Ville Hellman

Ville Hellman

Staff Engineer, Datadog
Ville is a Staff Engineer in Datadog's AI DevX group, where he builds evaluation infrastructure for the agents and tooling Datadog engineers use every day. He writes on AI-augmented engineering, AI literacy, and developer experience to make what's coming next easier for others to... Read More →
Friday September 18, 2026 13:15 - 13:40 CEST
G104 + G105 (Level 1)

16:20 CEST

Infrastructure Red Teaming With Abliterated Models: What Actually Stops Agent Attacks - Roy Belio, Red Hat
Friday September 18, 2026 16:20 - 16:45 CEST
Safety-aligned models refuse adversarial prompts, so you can't test whether your infrastructure controls actually work, but the models are all still susceptible to jail-breaking.
I removed that variable with an abliterated Qwen3.5 model to get zero refusals and 100% cooperation. Ran full suite of prompts with custom garak probes across three hardening tiers on an OpenClaw agent running in OpenShift.

I found out what worked and what gave false sense of security.
Sandbox isolation dropped credential exfiltration entirely in one step. NetworkPolicy killed cluster escalation. The prompt injection classifier caught encoding-based attacks. Three of four attack categories were fully stopped by Tier 2 (injection classification+isolation).

Memory poisoning was the exception. Probes that instruct the agent to write attacker content into its own memory continued to succeed across all tiers. OWASP added this as ASI06 to its 2026 Agentic Top 10. No deployed control addresses it today.

I'll present the full probe results, the defense configurations, and the open problem current agent architectures don't solve.
Speakers
avatar for Roy Belio

Roy Belio

Senior Software Engineer, Red Hat
Roy Belio is an AI Engineer at Red Hat, where he builds and evaluates proof-of-concept projects, drives open source contributions, and deploys AI/ML infrastructure on OpenShift.
Before Red Hat, Roy spent five years at Microsoft, Infinidat and Checkpoint.
He holds a BSc in Inform... Read More →
Friday September 18, 2026 16:20 - 16:45 CEST
G102 + G103 (Level 1)
 
Share Modal

Share this link via

Or copy link

Filter sessions
Apply filters to sessions.