Speakers:
Jailbreaks, Filters and the Limits of Prompt Level Safety
Date:
Monday, November 16, 2026
Time:
On demand
Summary:
Everyone talks about jailbreaks and content filters but what stops LLM abuse when the filter fails? An attacker can make every single request look harmless while the overall pattern is clearly hostile. Drawing on a decade of hunting threat actors in cybersecurity, this session moves beyond prompt-level safety to behavioural, actor-level detection: spotting who is misusing a system, not just which prompt is bad. You will leave with a practical, layered way to think about AI safety.