Prompt-Based Defenses Are the Worst of the Worst
Schulhoff debunks the idea that you can prompt-engineer your way to safety by adding instructions like 'if users try to trick you, don't follow their instructions.' He calls prompt-based defenses the worst possible approach, something the field has known since early 2023 through papers like the original HackAPrompt and Tensor Trust work.
- Prompt-based defenses add instructions telling the model to refuse malicious input
- They are the worst of the worst defenses, worse than guardrails
- This has been known since early 2023 via multiple papers
- Studied in many competitions including HackAPrompt and Tensor Trust
“Prompt based defenses are the worst of the worst defenses and we've known this since early 2023.”