Guardrails at 28ms: how we built a proxy that doesn't slow you down
A guardrail that adds a second of latency gets turned off. Here's the architecture behind Probe Guard's p95 under 30ms, and the trade-offs we made to get there.
Our first prototype of Probe Guard used a large language model to judge every request. It was accurate. It also added 900 milliseconds to every call, and the first design partner turned it off within a week.
That taught us the most important rule of guardrails: if it's slow, it doesn't exist. Here's how we got from 900ms to a p50 of 18ms and a p95 under 30ms.
Start with a latency budget
We set a hard budget of 30ms at p95 for the full input-and-output check, and worked backwards. Every component gets a slice: TLS and routing (4ms), input classifiers (8ms), PII detection (5ms), output streaming checks (amortized), and headroom.
Anything that couldn't fit in its slice was either redesigned or moved off the hot path.
Small, specialized classifiers
We don't use a foundation model in the request path. Instead, Guard runs a set of small classifiers, each trained for one job: injection detection, jailbreak patterns, encoded payloads, toxicity. The largest is under 100M parameters and runs on CPU.
Small models are less general, but they're fast, cheap, predictable and easy to calibrate per customer. When a classifier is uncertain, Guard can escalate to a slower async check without blocking the user.
We'd rather run five fast specialists than one slow generalist.
Run everything in parallel
Input checks are independent, so they run concurrently. The request is held for the slowest check, not the sum of all of them. On a typical request, PII redaction and injection detection finish within a millisecond of each other.
input ──┬─ injection (6.1ms)
├─ jailbreak (5.4ms)
├─ encoding (0.8ms)
└─ pii/redact (4.9ms)
↓ max = 6.1ms
→ model
Streaming without waiting
Output checks are harder, because users expect tokens to stream immediately. Guard checks output in rolling windows: tokens are released to the client as soon as the window they belong to is cleared. If a window fails, Guard cuts the stream and substitutes a safe completion.
The user-visible cost is roughly one window of delay, about 40 tokens, which is usually imperceptible.
Trade-offs we accepted
- Recall on novel attacks is lower than a big judge model would give. We compensate with weekly retraining from red-team output and async escalation.
- Grounding checks are lighter in the hot path. Deep factual verification runs asynchronously and reports to Trace.
- Per-customer tuning is required to hit the best precision. We think that's a feature, not a bug.
The numbers
On a three-node cluster at 5,000 requests per second: p50 overhead 18ms, p95 29ms, p99 41ms. False positive rate after two weeks of tuning averages 0.3% across customers.
Fast enough that, as one customer put it, "our users have no idea it's there". Which is exactly the point.