For operating companies
What it takes to run on AI.
Your engineers adopted AI on their own and output went up. The teams around them did not change, the review process did not change, and no record was kept of what any of it produced.
A governed AI operating model.
Where to start
Four questions every company running on AI has to answer.
Whoever you work with, and whether or not that is us. They are in the order the work has to happen. Not one of them is a question about tooling.
What do we build first, and who decides?
Agents remove the limit on how much can be built. Intent becomes the limit instead.
What stops an agent from doing something it shouldn't?
Today the answer is usually a system prompt and good intentions.
How do we know the work is right without reading all of it?
Review capacity is a ceiling. Writing the evals that lift it is a skill, not a purchase.
What has to be true before this runs past one team?
One team proving something is not the same as the company knowing it.
The answers
All four apply to one workflow at a time.
Scoped to a single workflow, each question has a plain answer and a small first step. Scoped to a whole company, none of them does.
Intent
A person owns the intent. The machine proposes the work.
A person writes what the workflow is supposed to do and approves it before anything is generated. Generated work then arrives as a candidate rather than as authority.
Control
Written where a person can read it. Enforced where the agent cannot reach it.
An agent with credentials can reach whatever those credentials reach. A system prompt is a request, and a request is not a control.
Verification
A proof shows the work met a standard set in advance.
An eval is the test. A proof is the retained record that this run passed it, kept where someone can find it a year later without asking the team that built it. When output rises faster than the capacity to review it, approval quietly turns into rubber-stamping.
Scale
The second team inherits whatever was written down.
If nothing was written down, the second team starts over and the company pays for the same lesson twice. The list of what has to be true is short and finite.
The unit of work
One workflow, with five things named.
Naming these five is what makes a workflow governable. Everything else gets built to serve it, and if the workflow can run without a component, the component waits.
The slice
Narrow enough to finish. A whole workflow is too big. One decision inside it is not.
The owner
A named person. Not a committee, and not a role that nobody currently holds.
The baseline
What it costs and how long it takes today, measured before anything changes.
The verifier
What correct means for this work, and who is entitled to say so.
The boundary
Where the agent stops and a person picks the work up.
Then a proof
Two figures come out of the record. Cost per verified outcome, and execution yield. Both compare with the next workflow.
What it asks of the company
Not a tooling decision.
Very little of this is a purchase. Most of it is a change to how work is specified and checked.
What does not change
- Business data stays where it lives
- Systems of record survive, and become inputs
- Nobody is asked to rip anything out to begin
- The first workflow runs alongside the current one
What does change
- Who specifies work and who checks it
- How a team is measured, from output to verified output
- Where the operating infrastructure sits, and who owns it
- What a manager reviews, and how long it takes them
Below a certain size, no one is accountable for the operating infrastructure inside the company. AI then defaults to the product organization, where it is measured on what ships, and the inside of the business is starved.
Who signs off
The security review is not a later step.
Three questions decide it, and they are the same three every time. What identity does this agent hold. What can it reach while holding it. What happens when something it reads tells it to do something else.
Bring the security owner to the conversation where the first workflow gets named. Not to the rollout.
Where this starts
Name one workflow.
Not a strategy, and not a platform decision. The first verified workflow has to meet three tests, and it gets chosen in a conversation rather than in a proposal.
It matters
Something the business would notice if it got better. A toy problem proves nothing worth having.
It repeats
Often enough that the second run teaches you something the first one could not.
It can be checked
Correct is decidable. If nobody can say whether an output is right, nothing can be proved about it.
Why this compounds
The second workflow costs less than the first.
A company that improves itself is not something anyone buys, and it is not a destination on a roadmap. It is what a business becomes once every verified result makes the next one cheaper to produce.
Name one and we will tell you what it would take to run it under a record.