Methodology

Graded on evidence, not vibes

Five phases turn 'is our AI safe?' into a reproducible, auditable answer.

01

Scope & threat model

We map every tool, data source and action your AI system can reach. The output is a permission graph: who can trigger what, under which conditions. This graph defines the test surface — nothing is tested ad hoc.

Permission graph & test plan

02

Test-case generation

From the permission graph we generate hundreds of adversarial scenarios: direct injection, indirect injection via documents and emails, multi-step goal hijacking, and cross-tenant access attempts. Cases are versioned so results are comparable over time.

Versioned test suite (200–500+ cases)

03

Controlled execution

Tests run against a staging or sandboxed environment with full logging. Every prompt, tool call and response is captured. We never test against production data, and destructive actions are always simulated, never executed.

Complete execution traces

04

Evidence grading

Each result is graded on evidence, not model judgement. A test fails only when the trace shows the unauthorized action or disclosure actually occurred or would have occurred. Borderline cases are reviewed by a human analyst.

Graded results: pass / fail / inconclusive

05

Report & remediation

You receive a report that ranks findings by exploitability and blast radius, with the exact trace that proves each one, plus concrete remediation guidance. Re-tests of fixed findings are included in every plan.

Ranked report + free re-test

Principles we don't compromise on

Evidence over opinion

A finding exists only if the execution trace proves it. We never report 'the model might…'.

Reproducibility

Every test case is deterministic and versioned. Run it again, get the same verdict.

Attacker's view

We test from outside, with no source-code access — the same position a real attacker starts from.

No production risk

All testing happens in isolated environments. Destructive actions are simulated, never executed.

Human review

Automated grading is fast; human analysts resolve the ambiguous cases that automation gets wrong.

Continuous, not annual

Models, prompts and tools change weekly. Permission testing must run as often as you ship.

See the methodology applied to your system

The free audit runs phases 1–4 on one application and shows you exactly where you stand.

Request a free audit