LLenny's Podcast
← All frameworks
StrategySander Schulhoff

The AI Security Vendor Due-Diligence Test

Five questions that expose whether an AI guardrail vendor is selling real protection or theater

Difficulty
Easy
Time to result
~days to results
Steps
5
Confidence
88%

Schulhoff argues most AI guardrail and automated-red-teaming vendors sell false confidence: automated red teaming always finds attacks (proving nothing), and guardrails don't meaningfully reduce an effectively infinite attack space. This is a set of pointed questions a buyer can ask to cut through aggressive marketing and fabricated statistics before purchasing.

Origin

Drawn from Sander Schulhoff's experience running AI red-teaming competitions that repeatedly broke commercial guardrails, plus a joint research paper with OpenAI, Google DeepMind, and Anthropic in which human attackers broke 100% of state-of-the-art defenses in 10-30 attempts.

Core principles

  • 01Automated red teaming always finds attacks against any transformer-based system, so a vendor's scary findings prove nothing novel.
  • 02A '99% catch rate' against an attack space of 1-followed-by-a-million-zeros is not statistically meaningful.
  • 03The only credible measurement is adaptive evaluation (attackers that learn), not static datasets.
  • 04Guardrails can create dangerous overconfidence in your security posture.

How to run it

  1. 1

    Ask them to red-team their own guardrail

    They ran their automated red-teamer against your models and found attacks. Ask what happens when they point it at their own guardrail. It will find plenty of attacks that work — because guardrails are themselves LLMs vulnerable to the same attacks.

    Pro tip Anyone can run this test; the vendor's reluctance is itself the answer.

  2. 2

    Ask how the 99% figure was measured

    Probe whether the effectiveness number comes from a static dataset of old attacks or from adaptive evaluation with attackers that learn over time. Static-dataset numbers are, in Schulhoff's words, quite useless.

    Pro tip Ask specifically whether human red-teamers were included — humans are still the strongest adaptive attackers.

    Watch out 99% of a near-infinite attack space still leaves basically infinite attacks; the percentage is marketing, not security.

  3. 3

    Test non-English and multilingual attacks

    Ask whether their models work on non-English languages. Many don't — which is disqualifying, because translating an attack into another language is a standard, trivial bypass.

    Pro tip If it only works in English, it's basically useless.

  4. 4

    Apply the 'smartest researchers' sanity check

    The best AI researchers at OpenAI, Google, and Anthropic have not solved adversarial robustness in years of trying. Ask why a vendor who doesn't even employ AI researchers would have.

    Pro tip Also check free/open-source alternatives — many are better than the paid products.

  5. 5

    Decide whether a guardrail buys you anything at all

    For a chatbot it adds no protection; for a determined attacker it's no obstacle; it doesn't even dissuade attackers. Conclude that in most cases the right number of guardrails to deploy is zero — and redirect the budget to permissioning and education.

    Pro tip Monitor and log inputs/outputs regardless — that's good deployment practice, not a security purchase.

    Watch out A guardrail's main effect may be making you overconfident, which is worse than deploying nothing.

In the wild

The competition that broke every guardrail

Schulhoff's team, alongside OpenAI, Google DeepMind, and Anthropic, threw adaptive attacks (RL, search-based, and human) at all state-of-the-art models and defenses including GPT-5. Human attackers broke 100% of defenses in roughly 10-30 attempts; automated systems needed orders of magnitude more tries and still only beat ~90%.

Empirically confirmed that commercial guardrails provide little real protection and that human adaptive attackers remain the true test.

Service Now Assist with protection enabled

Service Now's AI assistant had a prompt-injection protection feature switched on when a researcher demonstrated a second-order injection that recruited more powerful agents to perform database CRUD operations and send external emails.

The guardrail was enabled and the attacker still got through — a live example of protection-theater failing on a real system.

Common mistakes

Being impressed that a red-teamer 'found attacks on our models'

Automated red teaming finds attacks against every transformer-based model including the frontier labs', so a vendor showcasing your model's vulnerabilities has demonstrated nothing your competitors don't also have — it's a sales tactic, not a finding.

Treating a 99% ASR figure as meaningful security

With an attack space of 1-followed-by-a-million-zeros, catching 99% of a tiny tested sample leaves effectively infinite working attacks, so the headline number creates false confidence rather than real defense.

Is it for you?

Best for

CISOs and procurement leads evaluating AI guardrail or automated-red-teaming vendors

Not ideal for

Governance/compliance tooling and monitoring vendors, which Schulhoff considers genuinely useful and outside this critique

From the transcript

When these guardrail providers say we catch everything, that's a complete lie

00:00

what happens if they apply it to their own guardrail? Don't you think they'd find a lot of attacks that work? They would

the number of possible attacks is one followed by a million zeros

31:00

their models like like don't even work on non-English languages or something crazy like that which is ridiculous because translating your attack to a different…

36:00

the smartest artificial intelligence researchers in the world are working at frontier labs like OpenAI, Google, Anthropic, they can't solve this problem

36:30

From the episode

The coming AI security crisis (and what to do about it)

Sander Schulhoff