A security attack where malicious instructions are embedded in content that an AI is asked to process. If the AI follows those hidden instructions instead of (or in addition to) your request, it could take unintended actions. Most relevant when AI is processing untrusted external content — emails, web pages, documents from unknown sources. See the Prompt Injection guide.
Prompt Injection
Lessons that cover this
- What Happens to Your DataWhere it goes, who can see it, whether it trains the model — and the questions to ask before you upload anything.
- Writing an AI Policy People Will FollowOne page, five sections, one meeting. A twenty-page policy is the same as no policy.
- Agents: When AI Stops Answering and Starts DoingThe word is everywhere and means almost nothing. Here is the part that actually matters.
Related terms
This definition is part of a free, structured course on how AI actually works.
Start learning