Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 18, 2026 12:40 - 13:05 CEST
The AI ecosystem suffers from a critical engineering immaturity: deploying stochastic models via manual "vibe checks." Operating autonomous agents at enterprise scale requires abandoning ad-hoc observation for strict, distributed systems rigor. This session introduces a deterministic, CI/CD-native evaluation architecture for Agent Skills, shifting from indeterministic to reliable software execution.

By adhering to the formalized capability standards defined by agentskills.io, we will deconstruct the transition from subjective testing to hermetic, code-driven audits. Attendees will learn to engineer scenario matrices that enforce strict cognitive boundaries via negative testing—guaranteeing agents safely reject out-of-scope triggers. We will demonstrate isolating execution within sandboxed environments to capture pristine telemetry: deterministic tool-call structures, system exit codes, and exact token utilization.

Crucially, we address the anti-pattern of relying on "LLM-as-a-judge" for critical path assertions. Instead, we architect a framework grading system invariants via AST parsing and JSON Schema enforcement to achieve instantaneous, hallucination-immune evaluation.
Speakers
avatar for SHUVA JYOTI KAR

SHUVA JYOTI KAR

Principal Engineer, Palo Alto Networks
Shuva is a Principal Engineer at Palo Alto Networks building secure enterprise AI platforms. An author of two upcoming books: Engineering the Data Agent Control Plane (O'Reilly) and Agent Skills in Action (Manning), an open-source contributor and former OpenDaylight committer, his... Read More →
Friday September 18, 2026 12:40 - 13:05 CEST
G104 + G105 (Level 1)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link