Loading…
17-18 September | Amsterdam, Netherlands
View More Details & Registration

IMPORTANT NOTE: Timing of sessions and room locations are subject to change.
Friday September 18, 2026 16:20 - 16:45 CEST
Safety-aligned models refuse adversarial prompts, so you can't test whether your infrastructure controls actually work, but the models are all still susceptible to jail-breaking.
I removed that variable with an abliterated Qwen3.5 model to get zero refusals and 100% cooperation. Ran full suite of prompts with custom garak probes across three hardening tiers on an OpenClaw agent running in OpenShift.

I found out what worked and what gave false sense of security.
Sandbox isolation dropped credential exfiltration entirely in one step. NetworkPolicy killed cluster escalation. The prompt injection classifier caught encoding-based attacks. Three of four attack categories were fully stopped by Tier 2 (injection classification+isolation).

Memory poisoning was the exception. Probes that instruct the agent to write attacker content into its own memory continued to succeed across all tiers. OWASP added this as ASI06 to its 2026 Agentic Top 10. No deployed control addresses it today.

I'll present the full probe results, the defense configurations, and the open problem current agent architectures don't solve.
Speakers
avatar for Roy Belio

Roy Belio

Senior Software Engineer, Red Hat
Roy Belio is an AI Engineer at Red Hat, where he builds and evaluates proof-of-concept projects, drives open source contributions, and deploys AI/ML infrastructure on OpenShift.
Before Red Hat, Roy spent five years at Microsoft, Infinidat and Checkpoint.
He holds a BSc in Inform... Read More →
Friday September 18, 2026 16:20 - 16:45 CEST
G102 + G103 (Level 1)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link