Four tools. One very paranoid platform.
Red Team finds the holes, Evals stops regressions, Guard blocks attacks live, Trace proves it to your auditors. Use one, or run the full loop.
Automated adversaries that never get bored.
Connect an endpoint and Probe generates targeted attacks against your exact system prompt, tools and retrieval sources. Multi-turn, adaptive, and mapped to OWASP LLM Top 10 and MITRE ATLAS.
Jailbreaks, direct and indirect injection, encoding smuggling, persona hijacks, and many-shot attacks.
Poisoned tool outputs, malicious documents in RAG, and unsafe function-call chains.
Failed attacks are rewritten and retried, the way a motivated human attacker would.
Severity-ranked findings, reproductions and fixes. PDF for the CISO, JSON for the backlog.
Unit tests for things that aren't deterministic.
Define what "good" means once, then run it on every commit, prompt change and model upgrade. Calibrated graders, golden datasets, and diffs you can actually read.
GitHub Actions, GitLab CI and Buildkite. Fail the build on a score threshold.
LLM and rule graders validated against human labels, with agreement scores shown.
Compare any two models or prompts side by side across thousands of cases.
Turn real production failures into permanent regression tests with one click.
A bouncer that adds 18 milliseconds.
An OpenAI- and Anthropic-compatible proxy that inspects every input and output against your policies. Streaming-aware, fail-open or fail-closed, your choice.
Small, fast classifiers we tune on your traffic. No foundation model involved.
Emails, cards, IBANs, health identifiers and API keys, masked before they reach a model.
Rewrites answers that promise things your knowledge base doesn't say.
Managed cloud, your VPC via Helm, or fully air-gapped on your hardware.
Every verdict, on the record.
A searchable, tamper-evident log of every request Guard saw, every test Evals ran and every finding Red Team filed. Built for SOC 2, ISO 42001 and EU AI Act evidence.
Step through a full multi-turn conversation with each policy decision annotated.
Route spikes to Slack, PagerDuty or your SIEM within seconds.
Pre-mapped control packs for auditors. Fewer screenshots, fewer meetings.
Choose what is stored, where, and for how long. Hash-only mode available.
Sits in the path. Never in the way.
Guard is a stateless proxy you can scale horizontally. Your keys stay yours, and prompts never train anything.
p50 18ms, p95 29ms at 5k rps on a three-node cluster.
Zero retention mode. Nothing written to disk unless you turn on Trace.
Fail-open or fail-closed per route, with health checks and circuit breakers.
Bring your own provider keys. We never resell tokens or route to other models.
Block the merge, not the launch.
Add one step to CI. Probe runs your eval suite and a targeted red-team pass on every pull request, then comments the results inline.
name: probe on: [pull_request] jobs: safety: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: lanbolab/probe-action@v3 with: suite: evals/support-bot red_team: quick # 400 targeted attacks fail_under: 0.95 api_key: ${{ secrets.PROBE_KEY }}
The teams who get paged.
Platform & ML engineering
Ship model upgrades in days, not quarters, with diffs that show exactly what changed.
Security & AppSec
Continuous AI pentesting with findings mapped to the frameworks you already report on.
Risk & compliance
Evidence for SOC 2, ISO 42001, HIPAA and the EU AI Act, exported in a click.
Product leaders
A single safety score per release, so "is it ready?" has an answer.