Attempts to bypass an AI's safety guardrails through creative prompting — getting the model to say or do things it's designed to refuse. AI companies actively test for jailbreaks and update models to resist them. From an organizational standpoint, documented jailbreak attempts signal that AI policy enforcement is needed.
Jailbreaking
Related terms
This definition is part of a free, structured course on how AI actually works.
Start learning