The red-team playbook we give every new customer
The exact five-phase process our researchers follow in the first week of an engagement, so you can run a version of it yourself.
When a new customer signs up, the first thing we do is try to break their app. Not with the full 4,000-recipe library on day one, but with a structured process that finds the worst problems fast. Here it is, so you can run it yourself.
Phase 1: Map the blast radius
Before attacking anything, list what the system can actually do. What data can it read? What tools can it call? Who sees its output? A chatbot that only answers FAQ questions has a very different risk profile from an agent that can issue refunds.
We rank every capability by impact: read (can leak data), write (can change state), and speak (can say something damaging in your brand's voice).
Phase 2: Baseline attacks
Run the classics first. They still work more often than you'd think.
- Direct instruction override ("ignore previous instructions")
- System prompt extraction
- Role-play and persona attacks
- Requests for out-of-scope or harmful content
- Simple encoding: base64, leetspeak, other languages
If more than a handful of these succeed, stop and fix them before going deeper. There's no point testing advanced attacks on a system that falls to basic ones.
Phase 3: Targeted attacks
Now use what you learned in Phase 1. If the app can read customer records, try to get it to read someone else's. If it can call a refund tool, try to make it refund more than it should. Attacks written for your specific capabilities find the findings that matter.
Generic jailbreaks make headlines. Targeted attacks make incidents.
Phase 4: Indirect and multi-turn
Plant instructions in every source the system reads: documents, tool results, web pages, emails. Then run multi-turn conversations that escalate slowly over five to ten turns. This is where most critical findings in agentic systems come from.
Phase 5: Report and regress
Every finding gets a severity, a reproduction, and a recommended fix. More importantly, every finding becomes a permanent regression test in your eval suite, so it can never quietly come back.
Then automate it
Running this by hand is a great way to learn your system. Running it every week by hand is a great way to burn out your security team. Probe Red Team runs all five phases continuously and files findings straight into your backlog.