The world is handing AI agents the keys to move money. CONTINENTIA proves an agent cannot act above the level it was trusted with, and gives you a certificate anyone can verify. CONTINENTIA is our standard. Twelve levels. We defined it.
Talk to us →CONTINENTIA grades an agent's authority — how far it is permitted to reach, and whether it can reach itself. It answers one question: if this agent is wrong, hijacked, or hostile — how far does the damage go? It is an ordinal ladder of blast radius. Every rung is strictly more consequential, and strictly less reversible, than the one below it.
Every agent action is graded 1 to 12 before it executes. You certify an agent to a level; anything that reaches higher is stopped and recorded. The ladder has two halves. CL 1-7 — the agent ACTS ON THE WORLD: bounded, reviewable, reversible; a human can undo it. CL 8-12 — the agent ACTS ON ITSELF: this is loss of control, and no public scale grades these rungs.
| Level | What the agent may do | Example |
|---|---|---|
| CL 1-7 · the agent acts on the world — bounded, reviewable, reversible | ||
| CL 1 | read-trusted — read a source we control and trust | Check a price · look up a balance |
| CL 2 | analyse — reason over data we do not control, and report. The first rung an attacker can reach. | Summarise a supplier PDF · read a web page |
| CL 3 | communicate — emit to third parties in your name | Send an email · message a customer |
| CL 4 | transact — move money or assets | A transfer · a swap · sign_tx · a payout |
| CL 5 | commit — bind the organisation, technically or legally | Deploy to prod · merge to main · sign a contract · change DNS |
| CL 6 | control-access — issue, rotate or grant keys, credentials and permissions ← the amplifier | Create a key · grant a role · OAuth consent · add a user |
| CL 7 | actuate — command physical systems (OT, ICS, robotics, buildings, vehicles) | Open a valve · move a robot · change a setpoint · unlock a door |
| ↓ CL 8-12 · the frontier — the agent acts on itself. No public scale grades these. ↓ | ||
| CL 8 | acquire-capability — obtain new tools at runtime. It grows a new hand. | Install an MCP server · register a tool · add a plugin |
| CL 9 | self-modify — rewrite its own goals, memory, code, prompt or mandate. It edits the cage from inside. | Rewrite its own system prompt · alter its mandate |
| CL 10 | replicate — spawn copies, or recruit other agents. Blast radius stops being finite. | Spawn an agent · clone itself · recruit an agent (A2A) |
| CL 11 | self-fund — obtain accounts, pass KYC, acquire money or compute. It becomes self-sustaining. | Open an account · rent its own compute |
| CL 12 | self-persist — resist shutdown, conceal itself, distribute itself, deceive its overseers | Disable the kill switch · leave a copy behind |
CL 1-7, the agent acts on the world. CL 8-12, it acts on itself.
Everything below that line is a risk you can review, reverse and audit. Everything above it is loss of control.
Look at the ladder again: control-access (CL 6) sits ABOVE transact (CL 4) and commit (CL 5). That is deliberate, and it is the intellectual core of the standard.
A single fraudulent payment is bounded by the payment. A single fraudulent key is not bounded at all. An agent that can issue a key can grant itself the ability to transact — indefinitely, to any amount — and grant itself the ability to commit. Therefore any capability that can confer another capability must rank above it. That is what makes the ladder monotonic in blast radius.
Most access models get this backwards, and rank "manage access" below "move money" because a payment feels more consequential. It is the more consequential act. Access is the more consequential capability.
Scoring is deterministic and highest-match-wins: an action mentioning both "read" and "transfer" is a CL 4, never a CL 1. And an action we cannot classify — unknown, empty, unrecognised — scores CL 12, the worst rung on the ladder, and is therefore blocked by every ceiling below 12. We fail toward "this might be the end of the world", never toward "this is probably fine." An unknown action is not a safe action. It is an ungraded one.
Each action is scored on the twelve-rung ladder before it runs. An agent hired for research (CL 2) that suddenly attempts a transfer (CL 4) is stopped — by level, before it executes. Most stacks only ask "does this agent have a mandate, yes or no?" We ask "is this action above the level it was trusted with?"
Get a signed "Contained to Level N" certificate. Customers, auditors and partners can verify it instantly — and tampering with the level breaks the signature. Every certificate is stamped with the ladder version, so a level can never be silently re-read against a different scale.
Every containment breach is shared as a privacy-safe signature across the HERD network. One attack on any customer immunises the rest — your raw data never leaves you.
CONTINENTIA is derived from nothing. We built it from first principles, from facts anyone can observe: an agent reads, it analyses, it communicates, it moves money, it binds the company, it holds the keys, it touches machines — and it can rewrite itself. We publish crosswalks to other people's scales as a courtesy, and to prove we know the landscape. We derive from none of them.
| Scale | What it actually grades | Our crosswalk |
|---|---|---|
| UK AI Security Institute Frontier AI Trends Report, 18 Dec 2025, §6.3, Figure 23 |
Five levels grading MCP tool servers — how much autonomy a tool GRANTS. It does not grade agents. It was an analytical device used once, in a finance case study; it is not a standard, and no government uses it as one. | L1 → CL 1 · L2 → CL 2 · L3 "execution enabler" → CL 6 · L4 → CL 4, CL 5 · L5 "unrestricted execution" → CL 7 and everything above it, because their scale stops there. Stated, not hidden: CL 3 communicate has no clean counterpart on their scale. We map it to L2 as the least-bad fit and we say so. |
| Knight First Amendment Institute, Columbia Feng, McDonald & Zhang, 28 Jul 2025 |
Five levels grading the human's role — Operator through to Observer. How much of the human stays in the loop. | A different axis entirely. Complementary, not competing — it grades the human, we grade the agent. |
"AISI's five levels grade how much power a TOOL grants an agent.
Our twelve grade how far an AGENT can go — including the five rungs beyond execution,
where it begins to act on itself."
Knight grades the human's role. · AISI grades the tool's affordance. ·
CONTINENTIA grades the agent's authority.
You need all three — and only one of them goes past "unrestricted execution".
We publish our open rungs, exactly as we publish the open doors of our harness registry. Candidates under review: mass influence and persuasion at scale · training or fine-tuning new models · acquiring physical infrastructure. A standard that hides its gaps is marketing, not a standard.
We are not going to tell you that agents are getting good at opening bank accounts. The opposite is what the evidence says: on the UK AI Security Institute's RepliBench evaluation, current models fail the Know-Your-Customer check. KYC is currently holding. CL 11 self-fund exists for the day it does not. That is the honest position, and it is the stronger one — the rung is built before the capability arrives, not after.
Paste a CONTINENTIA certificate to check it's genuine, what level it certifies, and whether it's still valid.
Verification is open — anyone holding a certificate can confirm it. A certificate a stranger cannot check is a marketing claim, not a standard.
Sources · UK AI Security Institute, Frontier AI Trends Report, 18 December 2025, §6.3, Figure 23. (The institute was renamed from "AI Safety Institute" to "AI Security Institute" on 14 February 2025.) · UK AI Security Institute, RepliBench. · "Levels of Autonomy for AI Agents" — Feng, McDonald & Zhang, Knight First Amendment Institute at Columbia, 28 July 2025. · The crosswalks above are ours. Neither institution has produced or endorsed them.