Skip to content
Signal · August 27, 2026

The Competitive Challenges Facing Our Clients

Three immediate challenges, and a fourth that may prove the most consequential.

We see three immediate challenges facing the companies we advise and a fourth, longer-term threat that may prove the most consequential.

1. Ungoverned agent deployment creates a self-inflicted enterprise risk

Deploying agents broadly without governance, security, accountability, or management infrastructure is effectively the same as hiring 50 people, giving them access to company systems and data, and telling them to do whatever they want.

Many companies are doing some version of this deployment today. They are distributing agent capabilities faster than they can establish identity, permissions, trust boundaries, approval requirements, containment limits, escalation paths, or evidence of what those agents are actually doing. We judge that a company operating this way has roughly six months from the point its agents reach production scale before the approach produces a material failure of its own.

The consequences could include PII leakage, security breaches, regulatory violations, corrupted systems of record, uncontrolled external actions, and significant brand impairment. In a severe case a single incident could dramatically reduce a company's valuation or threaten its continued existence. The quiet accumulation of many invisible incidents likely carries the same weight.

Uber, among the most sophisticated engineering shops in the market, reportedly exhausted its entire 2026 AI budget by April as agent and coding-tool token consumption outran any measure of shipped outcome. To build the governance the initial rollout had skipped, its engineers stood up an agent registry, cryptographic agent identity, and gateway enforcement in front of internal systems. Everything short of these measures proved inadequate. Our architecture carries all of them alongside several further controls.

From here the exposure widens, because the agents are learning to run longer. METR, the evaluation group that measures how long a task a frontier model can finish unattended, has watched that horizon double roughly every seven months for six years and roughly every three since 2024. The models it graded this spring cleared sixteen-hour tasks at coin-flip reliability. METR now warns that its own suite can no longer measure the frontier. The successor family shipped in June with no measurable horizon at all. We understand at least two more sit trained behind it, awaiting clearance to release. Each doubling pushes more work into the background, where proactive agents act for hours and soon days with nobody watching the middle. A company that keeps delegating on that curve is flying progressively blinder into its own operations. A board asked to sign off on days of unobserved autonomous action will recognize the exposure without any help.

Line chart of the task length frontier models finish unattended at 50 percent success, doubling roughly every seven months since 2019 and every three months since 2024, with the June 2026 Claude 5 family released with no measurable horizon.
Figure 1. The autonomous task horizon, from METR estimates. Filled points carry measurements and the open marker shipped without one.

The answer is to make governed execution the prerequisite for scale, not to slow adoption. Every agent must have an identity, an accountable owner, bounded authority, approved data and execution surfaces, explicit escalation rules, and a complete evidence trail. Candidate workflows earn execution by meeting these conditions and earn the scale decision when the evidence shows the work safe and the economics sound.

2. Traditional competitors will use agents to create a widening execution advantage

While some companies struggle to govern AI, others will deploy it successfully across product development and operations. They will shorten time to market, increase engineering velocity, generate and test more code, release more features, and improve their products at a rate that was previously impossible.

That speed differential creates a straightforward competitive problem, because established companies now compete against the rate at which their rivals learn, build, ship, and improve. A market-leading product buys less protection every successive quarter.

A company that takes twelve months to deliver what an AI-enabled competitor ships in three is on a path toward market irrelevance. The gap compounds because every faster cycle returns more customer feedback, more operating evidence, and more reusable knowledge for the next feature, capability, or patch.

To the above point, Shein compresses design-to-shelf to roughly a week, launches thousands of small test batches every day, and scales only what sells, while legacy retailers like Gap still plan seasons six to nine months out. That gap ran on data and feedback rather than any single product advantage and it re-sorted a global category before agents existed. Agents now bring the same cycle to software and company operations.

3. AI-native startups will attack incumbent economics through the 10–100–3 model

The third challenge comes from companies designed around what we call the 10–100–3 model, in which fewer than 10 people generate $100 million in annual recurring revenue within three years of inception.

Cursor crossed the $100 million revenue line twenty months after launch with a team of a few dozen and Lovable did the same inside eight months with 45 people. Neither ran a ten-person roster at the milestone. The trajectory matters more than the roster, because each generation of these companies ships with fewer people, better tooling, and stronger unit economics than its predecessor.

Companies like these are AI-native from the first commit. Agents sit inside every function, the org chart never grows the layers an incumbent carries, coordination costs stay low, and decision cycles run in days rather than quarters. The unit economics that fall out of that structure bear little resemblance to the economics of an established SaaS company.

When SaaS businesses deploy governed agents successfully, the effect lands on the pricing unit itself. Seat-based revenue assumes a human behind every license and agents break that assumption. When a customer's agents absorb the work of five licensed users, per-seat revenue compresses at renewal even while the product stays fully competitive. AI-native entrants sharpen the squeeze by pricing on usage and outcomes rather than on people. The erosion arrives through net revenue retention well before it appears in a lost deal, which makes that metric the leading indicator worth watching.

Purchasing the same AI tools as the entrants is not a response, so incumbents must redesign how work is defined, governed, executed, measured, and improved. Tool adoption without operating-model change will produce activity, but it will not produce an AI-native company.

4. Frontier-model providers represent a longer-term vertical threat

By our read, the fourth challenge arrives eighteen to thirty-six months out, because the frontier labs have entered a recursive cycle where increasingly capable models accelerate further model and product development. Both private discussions and the labs' own public posts point the same direction. Expect continued, potentially dramatic gains in model intelligence and capability from more than one lab at once.

In August, OpenAI published benchmark results for Jalapeño, its first custom inference chip, co-developed with Broadcom from design to tape-out in nine months with the company's own models accelerating parts of the design work. Sam Altman announced it in eight words, "we made a chip and it is fast." The company's release goes further and says the models being served to users are now improving the infrastructure that will run the models after them. Nvidia proved the method four years ago, when nearly thirteen thousand AI-designed circuit instances shipped inside its Hopper generation. The cycle the labs describe in public now ships as hardware.

As intelligence becomes commoditized, the frontier providers will move further into the application layer. Their conduct to date settles any question of neutrality. They already build products, target workflows, and enter the industries they find attractive. Their path runs through the markets where our clients create value and make money. For some clients, the arrival could threaten the franchise itself.

In January, OpenAI and Anthropic entered healthcare within four days of each other, shipping consumer health products and enterprise suites into one of the most regulated verticals on earth. In April, Anthropic launched Claude Design into Figma's market. Figma's stock touched annual lows around the launch, its gross margin had already fallen six points on rented-inference costs, and one of its model suppliers now sells against it. In May, Claude for Legal arrived with twelve practice-area plugins and sold directly against Harvey, an eleven-billion-dollar platform that runs partly on Claude's own models.

The two of us met at Intel and spent years there under Andy Grove, who liked to reach for the creosote tree. The desert plant poisons the ground beneath its own canopy so nothing grows in its shade. He used it for the climb Intel never managed, from the platform up into the layers above it. The frontier labs have partially solved the problem Grove never could, which leaves every company renting their intelligence growing in creosote shade.

In pharma, Alphabet's Isomorphic Labs put an AI-designed drug into FDA-cleared human trials in January and raised over two billion dollars in May with Novartis, Lilly, and Johnson & Johnson signed as partners. In recruiting, OpenAI is standing up a jobs platform aimed straight at LinkedIn while naming education and finance as the markets it wants next.

Underneath those launches sits a leading signal, because model intelligence is approaching the median knowledge worker's and every lab knows enterprise adoption is the bottleneck that remains. Each lab now runs what amounts to a fast-growing army of knowledge workers in a datacenter. The incentive that follows is to work through the economy's profit pools one at a time and substitute model labor wherever a buyer will allow it.

Model reasoning scores plotted against the human IQ distribution, ranging from Claude 4.8 Opus at 117 to Grok 4.5 High and GPT-5.6 LUNA Max at 147, with the top two percent of humans marked at 130.
Figure 2. Model reasoning scores against the human IQ distribution, from TrackingAI. Open markers are the public Mensa Norway test and filled markers are leak-resistant offline tests built from unpublished puzzles.

The entries into healthcare, pharma, design, legal, and recruiting show the first form, where a lab judges the profit pool worth taking and competes against every incumbent in it. Finance looks next, because there is too much money in it and certified financial planners we know have already been asked by the labs to help train the coming models. Application capital competes with compute for every dollar, though, so no lab can run that assault everywhere at once. The second form appears where the criteria cut the other way, in smaller addressable markets, under rampant anti-competitive scrutiny, or behind network effects and domain knowledge no benchmark can verify.

In those markets a lab anoints rather than enters. It goes vertical by vertical and king-makes the biggest incumbent willing to move at model speed. Exclusive access arrives with forward-deployed engineers and shifts to revenue share once token budgets outgrow procurement. The chosen company runs roughshod over its competitors while the exclusivity wears the language of safety and anti-distillation. Suddenly the most capable model is too dangerous for broad API release. One king-made incumbent per vertical is enough, because every remaining board draws the intended conclusion.

Watching a rival get crowned moves a Fortune 500 board from deliberation to panic inside one meeting. Either play ends the same way for the rest of the vertical, so absorption capacity decides the outcome on both sides of the selection. The crowned incumbent has to metabolize capability arriving faster than any org chart adjusts. Everyone else has to outrun a rival whose head start is measured in model generations.

The strategic response is to build a durable layer of proprietary value above the models. The infrastructure worth owning is a domain and entity model, a governed execution system, a proprietary evaluation corpus, policy and routing memory, organizational memory, evidence of outcomes, and a reusable library of operating patterns. Static archives decay, so the data inside that layer differentiates only while it stays rights-secured and structured and keeps compounding through fresh sources and outcome feedback.

The models must remain replaceable, while the system that understands the company's domain, governs action, learns from execution, and compounds customer-specific knowledge must belong to the company.

What the four challenges demand

All four challenges run the same race between the pace a company can move and the control it can prove, a race decided at the operating model rather than the tool stack. Uber's answer was an agent registry and enforced identity. Shein's answer was a feedback system that turns demand signal into production faster than a planning calendar can. The labs' answer is to enter the vertical or crown its king, whichever the criteria favor. Ours is the same system in every case. Govern execution so speed is safe, prove outcomes so scale is earned, and build the owned layer so the models stay replaceable underneath it.

In practice the system arrives proof by proof. A readiness read locates the constraint and one governed workflow makes the method concrete. The proof package earns the scale decision and the assets it leaves behind lower the cost of the next.

A company that runs this cycle converts the four challenges from threats into a sorting mechanism that favors it. Speed without governance invites an existential operating failure, governance without speed forfeits the market, and the winners will be the companies that solve both at the same time.

Next step. Baser Potential works with software companies and their investors. Reach us at hello@baserpotential.com or www.baserpotential.com.