Skip to content

For operating companies

What it takes to run on AI.

Your engineers adopted AI on their own and output went up. The teams around them did not change, the review process did not change, and no record was kept of what any of it produced.


A governed AI operating model.

Where to start

Four questions every company running on AI has to answer.

Whoever you work with, and whether or not that is us. They are in the order the work has to happen. Not one of them is a question about tooling.

01

What do we build first, and who decides?

Agents remove the limit on how much can be built. Intent becomes the limit instead.

02

What stops an agent from doing something it shouldn't?

Today the answer is usually a system prompt and good intentions.

03

How do we know the work is right without reading all of it?

Review capacity is a ceiling. Writing the evals that lift it is a skill, not a purchase.

04

What has to be true before this runs past one team?

One team proving something is not the same as the company knowing it.

The answers

All four apply to one workflow at a time.

Scoped to a single workflow, each question has a plain answer and a small first step. Scoped to a whole company, none of them does.

Intent

A person owns the intent. The machine proposes the work.

A person writes what the workflow is supposed to do and approves it before anything is generated. Generated work then arrives as a candidate rather than as authority.

Control

Written where a person can read it. Enforced where the agent cannot reach it.

An agent with credentials can reach whatever those credentials reach. A system prompt is a request, and a request is not a control.

Verification

A proof shows the work met a standard set in advance.

An eval is the test. A proof is the retained record that this run passed it, kept where someone can find it a year later without asking the team that built it. When output rises faster than the capacity to review it, approval quietly turns into rubber-stamping.

Scale

The second team inherits whatever was written down.

If nothing was written down, the second team starts over and the company pays for the same lesson twice. The list of what has to be true is short and finite.

The unit of work

One workflow, with five things named.

Naming these five is what makes a workflow governable. Everything else gets built to serve it, and if the workflow can run without a component, the component waits.

The slice

Narrow enough to finish. A whole workflow is too big. One decision inside it is not.

The owner

A named person. Not a committee, and not a role that nobody currently holds.

The baseline

What it costs and how long it takes today, measured before anything changes.

The verifier

What correct means for this work, and who is entitled to say so.

The boundary

Where the agent stops and a person picks the work up.

Then a proof

Two figures come out of the record. Cost per verified outcome, and execution yield. Both compare with the next workflow.

What it asks of the company

Not a tooling decision.

Very little of this is a purchase. Most of it is a change to how work is specified and checked.

What does not change

  • Business data stays where it lives
  • Systems of record survive, and become inputs
  • Nobody is asked to rip anything out to begin
  • The first workflow runs alongside the current one

What does change

  • Who specifies work and who checks it
  • How a team is measured, from output to verified output
  • Where the operating infrastructure sits, and who owns it
  • What a manager reviews, and how long it takes them
This changes how the company works, so it needs a sponsor who can change how the company works.

Below a certain size, no one is accountable for the operating infrastructure inside the company. AI then defaults to the product organization, where it is measured on what ships, and the inside of the business is starved.

Who signs off

The security review is not a later step.

Three questions decide it, and they are the same three every time. What identity does this agent hold. What can it reach while holding it. What happens when something it reads tells it to do something else.

Answered before the build, the review is short. Answered after, the program stalls at exactly the point it was supposed to scale, and the stall gets attributed to the technology rather than to the sequence.

Bring the security owner to the conversation where the first workflow gets named. Not to the rollout.

Where this starts

Name one workflow.

Not a strategy, and not a platform decision. The first verified workflow has to meet three tests, and it gets chosen in a conversation rather than in a proposal.

It matters

Something the business would notice if it got better. A toy problem proves nothing worth having.

It repeats

Often enough that the second run teaches you something the first one could not.

It can be checked

Correct is decidable. If nobody can say whether an output is right, nothing can be proved about it.

Why this compounds

The second workflow costs less than the first.

A company that improves itself is not something anyone buys, and it is not a destination on a roadmap. It is what a business becomes once every verified result makes the next one cheaper to produce.

Name one and we will tell you what it would take to run it under a record.