Agentic Levels

The five levels of an agentic enterprise

Levels one to five did for self-driving what the industry has lacked for enterprise AI: a scale diagnosed from evidence, that an architect can argue with, and that tells a CEO what the next rung is worth.

Why a scale

“Make the company AI” has no definition.

A $20 chat subscription, an MCP server bolted onto the product and a firewall around everything have all been called that. The levels replace the slogan with a measurement: what can an agent do against your systems today, and what would it take to do more.

A human got a screen, a login and judgement. An agent needs a tool, an identity and a guardrail.
Apps were built for humans. Now agents want to use them.
The ladder

Five rungs, each diagnosed from evidence.

The ladder from L0, no callable interface, to L5, autonomously negotiate, with writes starting at L2.
The use case sets the level. The level is diagnosed from app, data and infrastructure evidence.
L0 · UI only, no APIs
No interface an agent can call. Where most applications start, and where 29 of 41 sat at the large US retailer we diagnosed.
L1
Headless read
Read and search
What an agent can do
Read your product. Check a balance, a price, a stock level.
Maps to
An MCP resource.
Public benchmark
Macy's (published): a concierge over 2.5M SKUs, about 4.75× revenue per visit, no application touched.
What it unlocks
Your product inside every customer's assistant.
L2
Transact
Click and buy
What an agent can do
Act within guardrails. Place the order, start the return.
Maps to
An MCP tool with per-agent identity and rate limits.
Public benchmark
Siemens (published): through an API it already had. Salesforce took fourteen months to climb from L1.
What it unlocks
Revenue that closes without a person in the loop.
L3
One workflow
Automate routine tasks
What an agent can do
Run a single team's workflow 80 to 95% end to end. Process the return, start to finish.
Maps to
MCP subscriptions and durable tasks.
Public benchmark
United Airlines (published): reasoning over a mainframe it never refactored, coverage grown from 15% to 50% of flights.
What it unlocks
A team's throughput without a team's headcount curve.
L4
Across the company
Orchestrate across systems
What an agent can do
Run multi-team, multi-system workflows on agents that authenticate to each other.
Maps to
The A2A agent card, inside your company.
Public benchmark
No public enterprise yet.
What it unlocks
The operating model itself.
L5
Across company borders
Autonomously negotiate
What an agent can do
Discover, trust and negotiate with suppliers', shippers' and regulators' agents.
Maps to
The A2A agent card, across companies.
Public benchmark
Nobody yet.
What it unlocks
The supply chain running itself.
How a level is diagnosed

Ten dimensions, scored from evidence, with no questionnaire.

Ten dimensions are scored per application from evidence, from a callable interface to per-agent identity. A use case's level is the lowest rung that every system it touches can support. That is why most enterprises diagnose at L0 or L1 even with a dozen pilots running.

Ten dimensions, D1 to D10, with the level each is required from and whether it is scored on the estate or the application.
Ten dimensions, scored from evidence. Five are estate properties, fixed once for every use case; five are per application, and those start to bite at L2.
Three public cases at L1, L2 and L3.
Three public benchmarks, from published accounts, placed on the ladder. One reads, one proposes, one executes, and none of them rewrote the system underneath. They are not ModernZ engagements.

Where is your enterprise?

Five weeks. Read-only. Free through cloud partner programs.