Runtime + Shadow AI · Free trial

A detection rate on its own
is worthless.

Anything scores a perfect 1.000 by blocking everything. So we measured 8 systems across 77 datasets and published the second column — what each one costs on legitimate traffic. Here is ours. Run it on your own stack with 10,000 free tokens.

Benchmark for Runtime Security · 8 systems · 77 datasets
0.9008
Attacks blocked
Across all 77 datasets
0.9185
Legitimate work delivered
The column nobody prints
+35pts
Detection gain from origin
0.617 → 0.967, InjecAgent
57–388ms
Scan time scales with input
Every window classified

Every figure carries its method and raw runs, including the four where we come second. *10,000 tokens per organization — no card, no procurement, no commitment.

See the console

The evidence, where you'll actually read it.

Verdicts, spans and the audit trail as they appear in the Runtime Security Console. Click through it yourself — no signup to watch.

Live traffic, one verdict column

Allow, flag, redact and block as they land — with the span that triggered each one.

Interactive walkthrough · dropping in
Where a CISO starts: what happened in the last hour.Recording 1 of 3
What's included

Both engines. One runtime.

Shadow AI covers what leaves the laptop. Runtime covers what reaches the model, and what it sends back.
Together they answer the question your board is already asking: what is our AI actually doing?

Shadow AIClient-side · discover & redact

Your team is already using AI tools you never approved. Shadow AI runs on the machine, names every one of them, and masks sensitive values before a model ever sees them.

  • Watches the reply, not just the prompt — the leak happens on the way out
  • Machine credentials, not only personal data: API keys and tokens in model output
  • Redaction happens on-device. The data never reaches a model, or us
Runtime Security ProxyServer-side · inspect & enforce

Indirect injection isn't a model problem, it's a contract problem. "Summarize the 2020 climate report" is work when your user types it and an attack when it arrives inside a fetched page. Declare where each span came from and the same classifier goes from 0.617 to 0.967 — with no new false positives.

  • Origin-aware inspection: user turn, document and tool output are not the same input
  • Classifies every window of a long input — an instruction in paragraph nine still gets caught
  • Every allow, flag and block written to a tamper-evident audit trail
One platform, one policy set, one tamper-evident audit trail — yours to export.
Model claims decay in six months. An architecture claim doesn't.
Included in the trial
How it works

Three steps. No procurement cycle.

Nothing to re-architect, no committee to convene. Point your traffic at the proxy and read your own second column by the end of the day.

Deploy in your cloud

The trial runs in the cloud — your tenant or ours. On-prem and air-gapped deployments exist outside this program; ask on the call.

FAQ

The method, and the terms.

Because the first one alone means nothing. Two of the eight systems we measured catch 97% of attacks — and flag 79 to 85% of legitimate traffic with them. That's roughly sixteen real requests refused for every extra attack caught. A detection rate you can't price in false positives isn't a result, it's a number.

The API contract, not the model. A scanner that takes one string and returns one verdict has thrown away the deciding information before it starts — the same sentence is work from your user and an attack from a fetched page. Declare the origin of each span and detection moves from 0.617 to 0.967 on InjecAgent, with zero movement on the clean controls. Same classifier, same rows.

Harmful-content classification: 0.7925, and we say so. We're also slower on short inputs — 57ms to 388ms as input grows, because every window gets classified rather than the first N tokens. A flat latency curve across a hundredfold change in input length isn't speed, it's a scanner that stopped reading. We publish the losses with the mechanism behind each one.

10,000 tokens of inspection per organization, across both engines, for the length of the trial. No card, no invoice, no automatic conversion to a paid plan. Cross the cap and we talk about what's next — we don't bill you for it.

It's the volume we've found is enough to surface something real in a working environment. The cap is per organization and may be revised as the program runs. If it changes, participants hear it from us first, in writing.

Both, combined, with the same detection rules and the same audit trail as a paid deployment. The trial limits volume, not capability.

Install and go. Shadow AI is an agent on the machine; the Runtime proxy sits in front of the endpoint you already call. No model changes, no application rewrite, nothing to re-architect.

The trial is cloud-only: your own tenant or VPC, or a public-cloud instance managed by Blindsight. Shadow AI redacts on the device, so sensitive values never reach a model or reach us. On-prem and air-gapped deployments exist, but not inside this program.

Nothing automatic. You keep your audit trail export, you get the full benchmark report, and if you want to continue we scope it then. Walking away costs nothing and needs no notice.

Get started

See it. Stop it. Prove it.

Take the numbers into every other vendor call — then measure ours on your own traffic. Tell us what you're building and a founder replies within one business day.