Skip to content

For operating companies

What it takes to run on AI.

Your engineers adopted AI on their own and output went up. The teams around them did not change, the review process did not change, and no record was kept of what any of it produced.


A governed AI operating model.

Why now

Four competitive challenges of the agent era.

Agents arrived in most technology organizations the way cloud did a decade ago, bottom-up and tool by tool, ahead of any decision about how they should be run. Three of the four challenges below are already on the technology leader's desk. The fourth sits eighteen to thirty-six months out and may prove the most consequential.

Each one is a race between the pace a company can move and the control it can prove. Each is decided in the operating model rather than the tool stack, which puts the answer inside the CTO's and CIO's remit.

On the desk now

Ungoverned agent deployment is a self-inflicted enterprise risk.

Deploying agents broadly without identity, bounded permissions, or a record of what they did is the operating equivalent of hiring fifty people and handing them production credentials with no job description. A company running this way has months, not years, before the approach produces a failure of its own, whether one visible breach or the quiet accumulation of many invisible ones.

Uber spent its full 2026 AI budget by April and published the governance it had skipped by May: an agent registry, an identity for every agent, and a gateway mediating each call into internal systems. The answer is not to slow adoption; it is to make governed execution the prerequisite for scale.

On the desk now

Competitors will turn agents into a widening execution gap.

The competitive unit has moved from the feature to the speed of the learning cycle behind it. The gap compounds, because every faster cycle returns more feedback and more reusable knowledge for the cycle after it. A market-leading product buys less protection every quarter.

Retail ran the experiment first. Shein reads demand in days and reorders only what sells, while the traditional apparel chain still runs six to nine months from forecast to fulfillment. That cycle re-sorted a global category before agents existed. Agents now bring it to software delivery, where a twelve-month roadmap competes against a rival shipping the same scope in three, and to the internal operations the CIO owns.

On the desk now

AI-native entrants will attack incumbent economics.

A new class of company reaches one hundred million in revenue with a few dozen people, inside two years, AI-native from the first commit. Each generation ships leaner than the one before, with agents inside every function and decision cycles that run in days.

For an incumbent that sells software, the pressure lands on the seat. Seat-based revenue assumes a human behind every license. When a customer's agents absorb the work of five licensed users, per-seat revenue compresses at renewal while the product stays fully competitive. The erosion shows up in net revenue retention well before it appears in a lost deal. Incumbents cannot respond by purchasing the same tools. They have to redesign how work is defined, governed, executed, and measured.

Eighteen to thirty-six months out

Frontier-model providers are a longer-term vertical threat.

The frontier labs appear to have entered a cycle in which each model accelerates the next, so dramatic gains from more than one provider should be the planning assumption. As intelligence commoditizes, the providers move up into the application layer. In January two labs entered healthcare within a week of each other. In April one of them shipped a design product into the market of a company it supplies.

The layer worth owning is infrastructure: a domain and entity model, a governed execution system, a proprietary evaluation corpus, and the record of what worked. Models must remain replaceable. The system that understands the company's domain and learns from its own execution must belong to the company.

Both at once

Speed without governance creates operating risk. Governance without speed creates competitive risk. Organizations that solve both at once turn the four challenges into a sorting mechanism in their favor.

Solving both is a change to how work is specified, controlled, checked, and scaled. The four questions below are those four decisions, in the order the work has to happen.

Where to start

Four questions every company running on AI has to answer.

Whoever you work with, and whether or not that is us. They are in the order the work has to happen. Not one of them is a question about tooling.

01

What do we build first, and who decides?

Agents remove the limit on how much can be built. Intent becomes the limit instead.

02

What stops an agent from doing something it shouldn't?

Today the answer is usually a system prompt and good intentions.

03

How do we know the work is right without reading all of it?

Review capacity is a ceiling. Writing the evals that lift it is a skill, not a purchase.

04

What has to be true before this runs past one team?

One team proving something is not the same as the company knowing it.

The answers

All four apply to one workflow at a time.

Scoped to a single workflow, each question has a plain answer and a small first step. Scoped to a whole company, none of them does.

Intent

A person owns the intent. The machine proposes the work.

A person writes what the workflow is supposed to do and approves it before anything is generated. Generated work then arrives as a candidate rather than as authority.

Control

Written where a person can read it. Enforced where the agent cannot reach it.

An agent with credentials can reach whatever those credentials reach. A system prompt is a request, and a request is not a control.

Verification

A proof shows the work met a standard set in advance.

An eval is the test. A proof is the retained record that this run passed it, kept where someone can find it a year later without asking the team that built it. When output rises faster than the capacity to review it, approval quietly turns into rubber-stamping.

Scale

The second team inherits whatever was written down.

If nothing was written down, the second team starts over and the company pays for the same lesson twice. The list of what has to be true is short and finite.

The unit of work

One workflow, with five things named.

Naming these five is what makes a workflow governable. Everything else gets built to serve it, and if the workflow can run without a component, the component waits.

The slice

Narrow enough to finish. A whole workflow is too big. One decision inside it is not.

The owner

A named person. Not a committee, and not a role that nobody currently holds.

The baseline

What it costs and how long it takes today, measured before anything changes.

The verifier

What correct means for this work, and who is entitled to say so.

The boundary

Where the agent stops and a person picks the work up.

Then a proof

Two figures come out of the record. Cost per verified outcome, and execution yield. Both compare with the next workflow.

What it asks of the company

Not a tooling decision.

Very little of this is a purchase. Most of it is a change to how work is specified and checked.

What does not change

  • Business data stays where it lives
  • Systems of record survive, and become inputs
  • Nobody is asked to rip anything out to begin
  • The first workflow runs alongside the current one

What does change

  • Who specifies work and who checks it
  • How a team is measured, from output to verified output
  • Where the operating infrastructure sits, and who owns it
  • What a manager reviews, and how long it takes them
This changes how the company works, so it needs a sponsor who can change how the company works.

Below a certain size, no one is accountable for the operating infrastructure inside the company. AI then defaults to the product organization, where it is measured on what ships, and the inside of the business is starved.

Who signs off

The security review is not a later step.

Three questions decide it, and they are the same three every time. What identity does this agent hold. What can it reach while holding it. What happens when something it reads tells it to do something else.

Answered before the build, the review is short. Answered after, the program stalls at exactly the point it was supposed to scale, and the stall gets attributed to the technology rather than to the sequence.

Bring the security owner to the conversation where the first workflow gets named. Not to the rollout.

Where this starts

Name one workflow.

Not a strategy, and not a platform decision. The first verified workflow has to meet three tests, and it gets chosen in a conversation rather than in a proposal.

It matters

Something the business would notice if it got better. A toy problem proves nothing worth having.

It repeats

Often enough that the second run teaches you something the first one could not.

It can be checked

Correct is decidable. If nobody can say whether an output is right, nothing can be proved about it.

Why this compounds

The second workflow costs less than the first.

A company that improves itself is not something anyone buys, and it is not a destination on a roadmap. It is what a business becomes once every verified result makes the next one cheaper to produce.

Name one and we will tell you what it would take to run it under a record.