AI-Sentinel

Independent guardrail assurance

Prove your AI guardrails work. Every month, and every time they change.

AI-Sentinel tests the guardrails on your GenAI apps from outside. It reaches only the guardrails you name, through read-only access you can remove at any time. You get an alert when protection drops, and signed scorecards your auditors can check for themselves.

No code changes None of your traffic No production data

Prompt-injection catch rate

Trial · 9 Oct 2026
Run 1, baselinePrompt-attack filter at its strongest
152 of 27555.3% · range 49.4–61.0%
Gap
Run 2, after a settings changeFilter lowered in the cloud console
110 of 27540.0% · range 34.4–45.9%
GapRetested on its ownDrift alert sent
Run 3, after the fixFilter restored
152 of 27555.3% · range 49.4–61.0%
GapAlert closed
0%100%

Dashed line: AI-Sentinel's 80% criterion, not an industry standard. Shaded band: the 95% range.

$ openssl dgst -sha256 -verify sentinel-public.pem \ -signature scorecard.json.sig scorecard.json Verified OK

One guardrail in a trial account set up as a customer would. The vendor isn't named: any guardrail needs measuring.

The problem

Guardrails are switched on. Few teams can show how well they work.

Settings change quietly

Lowering a guardrail setting in a cloud console raises no error. In our trial, one change cut the catch rate from 152 to 110 of 275 test attacks, and the app looked exactly the same.

Defaults aren't measurements

Vendors document what their filters are for, not how well they catch attacks on your use case. At its strongest setting, the guardrail we tested caught 141 of 263 attacks from a public prompt-injection test set.

Auditors want evidence

When a customer or an auditor asks what your AI was meant to block, and whether it did, most teams have screenshots. Sentinel gives you signed results.

Method

How it works

Sentinel works alongside the guardrails you already run. Continuous testing never sits in your request path.

  1. Connect in minutes

    Deploy our template in your AWS account. It creates a read-only role limited to the guardrails you name, and a locked bucket for your scorecards. Delete the role and testing stops.

  2. Set the limits

    A signed testing authorisation records your consent, the guardrails in scope and a monthly spending limit. A full run is about 1,287 test prompts; on the guardrail we tested, that cost about $0.52.

  3. We test on a schedule and on every change

    A versioned set of 1,269 test prompts, attacks and harmless ones, plus your own examples. Sentinel runs it every month and again whenever a guardrail's settings change.

  4. You get alerts and evidence

    Email and signed webhook alerts when protection drops, sent only to the contacts you give us. Each scorecard is signed and written to your bucket, where you can verify it yourself.

Results · 9 October 2026

What we proved, end to end

In a trial account set up the way a customer would set it up, with our own account testing it from outside.

Caught a weakened guardrail

Lowering one setting cut the catch rate from 152 to 110 of 275 attacks. Sentinel noticed the change, retested on its own and sent an alert. After the fix, the next retest showed 152 of 275 again and the alert closed.

Measured one guardrail as it is

At its strongest setting, it caught 141 of 263 attacks from a public prompt-injection test set (54%, 95% range 48–60%), and blocked 11 of 399 ordinary harmless prompts (2.8%).

Stayed inside the customer's limits

A run over the signed spending limit was refused before it sent anything. Every Sentinel session appeared by name in the customer's audit trail, and after the customer deleted the role, the next run was refused.

Left evidence the customer could check

Every scorecard was signed and verified on the customer's side with openssl alone, with no Sentinel software involved.

These are results from one guardrail configuration in a trial, not a ranking of vendors. Your guardrails, settings and use case will give different numbers, which is why they need measuring.

Reporting rules

How we report numbers

An assurance product is only as good as the honesty of its figures. These rules apply to every scorecard, alert and report.

Every figure shows its counts

Rates read as 152 of 275 with a 95% range, so you can see how far a number can be trusted.

Small samples get no percentage

Under 30 samples, a result reads too few to judge.

Unmeasured means unmeasured

Anything we haven't measured is labelled that way and never filled in with a number. When we can't inspect a guardrail, its status is unknown.

Mappings are a guide

Findings map to the EU AI Act, NIST AI RMF and the OWASP Top 10 for LLM applications as a guide for your auditors. Sentinel doesn't certify compliance.

The criteria are ours, and labelled

We judge catch rates against 80% and false blocks against 5%. Those are AI-Sentinel's criteria, shown on every scorecard, and they aren't an industry standard.

Signing key

Check a scorecard yourself

Every scorecard is signed with AI-Sentinel's key. Check it against the key published here, not against a key that arrives with the scorecard: anyone can sign a file with a key of their own.

Our scorecard signing key

ECDSA P-256 with SHA-256, in use from October 2026. Download it only over HTTPS: assurance-signing-2026-10.pem.

Its fingerprint

SHA-256 of the key (DER):

5b0a21c720c39bf30ad7451525131897e9b8a023aec770948d78ef59ea1bda3d

$ curl -sO https://www.sentinel.cloudcognoscente.com/keys/assurance-signing-2026-10.pem $ openssl pkey -pubin -in assurance-signing-2026-10.pem -outform DER | openssl dgst -sha256 5b0a21c720c39bf30ad7451525131897e9b8a023aec770948d78ef59ea1bda3d $ openssl dgst -sha256 -verify assurance-signing-2026-10.pem \ -signature scorecard.json.sig scorecard.json Verified OK

If a key is ever retired, it stays listed here with its dates, so older scorecards can still be checked.

Coverage

What Sentinel tests

Today

  • Amazon Bedrock Guardrails
  • Guardrails behind a LiteLLM proxy

Planned for November 2026

  • Azure AI Content Safety
  • Google Model Armor

Optional runtime gateway

For teams that also want enforcement: one approved policy on every prompt, reply and tool call, across 7 model providers and 5 checkers. It adds 52 ms at p99 in the same region as the model.

Design partners

Run a pilot with us

We're working with two or three design partners on an 8–12 week pilot on one GenAI app, judged against success criteria you set. No production traffic or customer data is needed: Sentinel sends its own test prompts to your guardrails.

You get

  • An independent catch-rate report on the guardrails you run today
  • Alerts when protection drops, from day one
  • Signed scorecards your risk and audit teams can verify themselves
  • Direct influence on the roadmap, and direct support from the people building it

We ask

  • The guardrails behind one GenAI app, in production or staging
  • A champion (AI platform, security or risk lead) and an engineer
  • A 30-minute feedback call each week
  • A reference if the pilot succeeds; anonymised is fine

Where we are

AI-Sentinel is early. Continuous testing was proven end to end on 9 October 2026 in a trial account set up as a customer would, and there's no customer yet. An external penetration test comes before any real customer data, and SOC 2 comes later. We'd rather tell you that now than have you find out later.

To talk about a pilot, write to