Gil Allouche
← All writing
GlossaryRunning a company with AI

AI-first company means AI does the work by default

The term carries two different claims — one about a product, one about how a company runs. Here is the definition, the origin, the neighbour it gets confused with, and the test.

September 27, 2026·Gil Allouche·6 min read
A glossary entry. The definition sits at the top in one sentence, with the context that makes it useful underneath.

AI-first company, defined

AI-first company
An AI-first company is one whose default way of getting work done is a model or agent, with humans handling exceptions, approvals and the work that is deliberately kept human.

Two things fall out of that sentence. If a human does the work and AI helps, the company is not AI-first. If AI does the work and a human signs off, it is.

The label gets attached at two levels: the product, and the operating model. Separate claims about separate things. Almost every argument about whether a company is "really" AI-first is two people using the two meanings at each other.

Where the term came from, and the second meaning it picked up

April 2016, Google's founders' letter. Sundar Pichai described a shift "from mobile-first to AI-first," and the argument was about product design — the interface should be a model that understands intent, not a screen full of buttons. That is the original sense. It still dominates product marketing.

The second sense is about internal work rather than the product. It says: the way this company produces its output — support replies, pipeline research, code, invoices, QA — runs through models by default, and headcount is what you add when that fails. A payments company with no AI anywhere in the product can be AI-first in this sense. A company whose entire product is a chatbot can fail it completely, because every internal process still runs through a human with a spreadsheet.

AI-first (product)AI-first (operating model)
ClaimThe model is the interfaceThe model is the labour
Dates fromGoogle, April 2016Current usage, 2024–2025
EvidenceWhat the user touchesWhat the org chart and the run logs look like
Fails whenModel output is wrong in front of the customerModel output is wrong and a human has to redo it anyway

When you see the phrase with no qualifier in 2025, it means the operating model. That is the version being sold in board decks, and it is the one that changes hiring.

What it is not: AI-enabled

AI-enabled — also sold as AI-powered — means AI sits inside a process that a human still owns. The rep writes the email with a suggestion panel open. The analyst gets a summary and then does the analysis. Real value, different architecture. The human is still the default path; the model is an accessory to it.

AI-enabled keeps the human on the default path; AI-first puts the agent on itAI-enabledRequestHuman runs stepAI assistsOutputAI-firstRequestAgent runs stepOutputHuman: exceptions

Two neighbours get mistaken for it. AI-native is about founding conditions — built after the models existed, no legacy system to migrate. It is a date, not a behaviour, and an AI-native startup founded last year can still route every single task through a person. AI transformation is a program with a budget and an end date. AI-first is a claim about the steady state after that program finishes.

The distinction that matters operationally is who picks the next step. A fixed sequence with a model in one box is a workflow; a loop where the model chooses what to do next is an agent, and Anthropic's engineering write-up draws exactly that line. I've argued the same thing at length in an AI agent decides its own next step.

A worked example with real numbers

Klarna, 27 February 2024. The company published figures on an OpenAI-powered customer service assistant that had been live one month: 2.3 million conversations, two-thirds of all its customer service chats, work it described as the equivalent of 700 full-time agents. Average resolution time dropped from 11 minutes to under 2. Repeat inquiries fell 25%. It ran in 23 markets, 24/7, in more than 35 languages, and Klarna estimated a $40 million USD profit improvement for 2024.

That is the operating-model sense, cleanly. Klarna's product is payments. The AI did not become the interface to the product — it became the labour behind a function that used to be staffed. And the number that makes it a real example rather than a demo is two-thirds: the model is on the default path, and humans take the third that routes to them.

One caveat belongs with the number. It is the company's own press release, at month one. Sustaining a share like that is a different exercise from reaching it, and support is the easiest function in the building to start with — high volume, bounded scope, mistakes that are cheap to reverse. A support result is not evidence that the same company can put an agent on revenue decisions.

How to tell whether something qualifies

Five tests. Each one is answerable in a sentence, and vagueness is itself the answer.

  1. Default path. Name one workflow. What runs first — the agent or the person? If the agent only starts when a human triggers it and a human finishes it, that is AI-enabled.
  2. What gets added when volume doubles. Seats, or capacity? AI-first companies buy tokens and compute. AI-enabled companies buy headcount and give it better tools.
  3. Cost per run. OpenAI and Anthropic publish per-token pricing, so this is arithmetic, not a mystery. If nobody can tell you what one run costs, nobody is operating the thing — they are demoing it. The standard operating advice is to build the meter first.
  4. What happens when it's wrong at 3am. Unbounded loops are the most common production failure: the agent retries, re-plans, re-calls and burns budget with no stopping condition. The fix is a step cap and a spend ceiling enforced outside the model, not a prompt asking it to be careful. No owner and no cap means you have a pilot.
  5. Measured, not felt. METR ran 16 experienced open-source developers across 246 real tasks in early 2025. With AI tools allowed they were 19% slower — and they reported believing they had been about 20% faster. Self-reports about AI productivity point the wrong way, so an AI-first claim resting on "the team says it's faster" is resting on nothing.
Unbounded loops are the most common way an agent in production becomes an expensive incident. Cap steps and spend outside the model, in code you control.

My own position, since I run agents in production on my own projects and pay the bill: I would not hand an agent spend authority on a live budget today, and the parts of go-to-market that involve a promise to a customer should not be delegated at all — which parts I'd leave alone is a separate argument. That is not a hedge. Naming what stays human is the thing that makes the rest of the claim checkable, and anyone genuinely running this way can list those parts in under a minute.

Sources

  1. Klarna AI assistant handles two-thirds of customer service chats in first month — Worked example figures: 2.3M conversations, 700-agent equivalent, 11 min to under 2 min, $40M profit estimate, 35+ languages, 25% fewer repeat inquiries
  2. Anthropic — Building effective agents — Workflow vs agent distinction; stopping conditions and guardrails as engineering requirements
  3. OpenAI API pricing — Per-token pricing is published, so cost per run is calculable
  4. Anthropic pricing — Per-token pricing is published, so cost per run is calculable
  5. Google Founders' Letter 2016 (Sundar Pichai) — Origin and date of the 'mobile-first to AI-first' framing

Related