Speakers:

Jyoti Yadav

Jailbreaks, Filters and the Limits of Prompt Level Safety

Date:

Monday, November 16, 2026

Time:

On demand

Summary:

Everyone talks about jailbreaks and content filters but what stops LLM abuse when the filter fails? An attacker can make every single request look harmless while the overall pattern is clearly hostile. Drawing on a decade of hunting threat actors in cybersecurity, this session moves beyond prompt-level safety to behavioural, actor-level detection: spotting who is misusing a system, not just which prompt is bad. You will leave with a practical, layered way to think about AI safety.

Ready to attend?

Register now! Join your peers.

Register nowView agenda
Newsletter Knowledge is everything! Sign up for our newsletter to receive:
  • 10% off your first ticket!
  • insights, interviews, tips, news, and much more about Machine Learning Week Europe
  • price break reminders