Permission testing for AI apps & agents

TokenVeto

Does your agent stay inside the lines?

Turn written access rules into executable checks. Test a local copy of your app, and keep the evidence behind every verdict.

permission suite · local buildIllustrative retest
TENANT-03

Aster user reads Cobalt record

FAILEDPASSED
REVOKE-07

Suspended user's run writes file

FAILEDPASSED

Local execution / Paired allowed-use controls / Evidence-backed verdicts

Prompt injection ✦Cross-tenant reads ✦Goal hijacking ✦Tool-scope escape ✦Memory poisoning ✦Privilege escalation ✦Data exfiltration ✦Action-approval bypass ✦Prompt injection ✦Cross-tenant reads ✦Goal hijacking ✦Tool-scope escape ✦Memory poisoning ✦Privilege escalation ✦Data exfiltration ✦Action-approval bypass ✦

Why TokenVeto

What makes the difference

Runs against a copy, never the live app

Checks execute against a local copy of your build with full tracing. Destructive actions are simulated — proven, never performed.

See how it works →

Every verdict carries its trace

A finding exists only when the execution trace proves it. No score is computed from a model's own account of what it did.

Read the methodology →

Zero cross-tenant reads

One customer never reaches another customer's records — verified per tenant, per role and per session, on every single run.

Explore use cases →

Graded on backend records

Where the backend is visible, the verdict cites the record that was or was not written — not the agent's claim about it.

Read the FAQ →

The three lines we defend

Permissions are promises. We test them.

No cross-tenant reads

One customer can never read another customer's records — verified per tenant, per role, per session.

Only allowed actions

The agent can only take actions its user is permitted to take, even when cleverly asked otherwise.

Revoked means revoked

Once access is removed, it stays removed — even mid-run, mid-conversation, mid-plan.

2 / 14

pilot runs wrote a file after the user was suspended

$4.99M

average cost of a data breach (IBM)

54%

of organizations have no approach to limiting agent access (Gartner)

Coverage

What TokenVeto checks for you

Injection & hijacking

Direct and indirect prompt injection, goal hijacking and tool-scope escape, run as adversarial scenarios against your real tools.

Isolation & revocation

Tenant boundaries, role limits and revoked access, verified mid-run and mid-plan — the moments a suspension actually leaks.

Allowed-use controls

Every forbidden check runs beside a control that must still pass, so a fix that over-blocks a real workflow fails instead of passing quietly.

The permission graph

Every connection your agent can reach — mapped and watched.

Tools, data sources and actions become nodes in a live graph. TokenVeto walks every edge with adversarial tests, so "allowed" is a fact, not an assumption.

  • Each node is a real capability of your app
  • Each edge is a tested permission path
  • Blocked paths are proven with execution traces
AGENTTOOLSDATAACTIONSBLOCKED
example run · suite v3.2RUNNING
Injection resistance92%
Tenant isolation100%
Action scoping78%
Revocation enforcement64%
verdicts graded on evidence412 / 438

A living report

Watch your defenses being tested, edge by edge.

Every suite run produces a live breakdown per defense line — what held, what failed, and the trace that proves it. Re-runs after a fix are free and automatic.

Explore the methodology

Setup

Point it at a rule, not a prompt.

Write the boundary in plain language — “Aster users may read Aster records, never Cobalt’s” — and TokenVeto turns it into executable checks against a local copy of your app. Runs and evidence stay on your machine.

Surfaces under test

Evidence first

No claims. Evidence.

TokenVeto does not take an agent's word for anything. A verdict is written only when the execution trace proves the outcome — and where the backend is visible, it cites the record that was or was not created.

That is how the finding that started this project surfaced: a suspended user's run wrote a file in 2 of 14 cases. A backend record caught it — no report, no self-assessment.

From rules to verdicts

Four steps to an evidence-graded answer

01

Map the permission graph

Every tool, data source and action your agent can reach becomes a node. The graph defines exactly what 'allowed' means — before any test runs.

02

Generate adversarial tests

Hundreds of versioned scenarios: direct and indirect injection, goal hijacking, cross-tenant access. Each forbidden check is paired with an allowed-use control.

03

Execute in a sandbox

Tests run against a local copy of your app with full tracing. Destructive actions are simulated — proven, never performed.

04

Grade on evidence

A finding exists only when the execution trace proves it. Verdicts cite backend records, not model claims. Fixed findings get a free re-test.

Graded on evidence

What the app did — not what the agent said.

TokenVeto grades what each test user could actually reach and, where it can see the backend, what was recorded. Every forbidden check runs beside an allowed-use control, so a fix that over-blocks a real workflow fails instead of passing.

Read how it works →
rule.yaml
allow: refund only if backend.refunds contains the record
TRUTH-01 · agent claims refundno record
ALLOW-03 · own-tenant refundrecord found
REVOKE-07 · suspended userrevoked

Test the build you're about to ship.

Runs and evidence stay on your machine. Export JSON, HTML reports or SARIF 2.1.0.

Results apply to the tested build and data and are not a safety certification.