What an AI-first company is, and what it is not
The term's 2016 origin at Google, the two different things people mean by it, and three tests that separate AI-first from AI-enabled.
AI-first company
- AI-first company
- An AI-first company is one where AI systems are the default doer of routine work and humans handle the exceptions, not one where employees use AI tools to work faster.
Two words carry the load: default and doer. Default means the work routes to a model before it routes to a person — configuration, not intention. Doer means the system produces the output of record: the reply that ships, the invoice that posts, the campaign that goes live. Not a draft a human retypes.
Everything else in the term is decoration.
Where the term came from: Google, April 2016
Sundar Pichai used it in the Alphabet Founders' Letter published April 2016, describing a shift "from mobile first to an AI first world." The construction borrows from "mobile-first," a web design phrase in common use from roughly 2009: design for the constrained device first, then scale up, instead of shrinking a desktop layout.
The 2016 version was a product claim. Build the thing so the model is the interface. It said nothing about who inside the company did the work.
The operational meaning attached later, once tool-using agents could complete multi-step tasks without a human stitching the steps together. That drift is why the term needs disambiguation before it is usable in a sentence.
Two meanings. Label which one you mean
| Sense 1: AI-first product | Sense 2: AI-first operating model | |
|---|---|---|
| Claim | The model is the core of what we sell | The model does the internal work |
| Evidence | Model in the product's critical path | Work items assigned to agents by default |
| Fails when | Model quality drops | Nobody owns the exception queue |
| Headcount | Normal | Deliberately flat |
| Typical era | 2016 onward | 2023 onward |
Metadata.io was sense 1. The product ran on models, I hold 6 patents in AI-driven marketing, and we went from $0 to $15M ARR on $50M raised. The company that built it was a normal company: people did the work. Calling that "AI-first" in 2019 was accurate under sense 1 and would have been a lie under sense 2.
Someone searching this term today wants sense 2. Say which one you mean in the board deck and on the careers page — the two have almost no operational overlap.
What it is not: AI-enabled
AI-enabled means humans still own the work and use models to go faster. A support rep with a suggested-reply panel. A seller with a research copilot. An engineer with autocomplete. Output goes up, the assignment of work does not change, and if you switch the tool off on a Tuesday everything still ships — slower.
AI-first changes the assignment. Switch it off and the queue stops.
AI-enabled is not a rung on the ladder to AI-first. It is a different bet — buy tools, keep the org chart — and it is a defensible one. What is not defensible is making that bet and describing it as a transition.
Two other neighbours get confused with it. AI-native is a founding condition, not an operating one — a company started after late 2022 whose first architecture assumed a model in the loop. A 40-year-old insurer can become AI-first. It cannot become AI-native. Digital transformation is about moving paper processes into software; AI-first assumes the software already exists and asks who operates it.
A worked example with real numbers: Klarna, 2024
Klarna published the cleanest documented case of sense 2. In February 2024 the company said its OpenAI-powered assistant had, in its first month live, handled 2.3 million conversations — about two-thirds of its customer service chats — doing the work equivalent of 700 full-time agents across 23 markets and 35 languages. Average resolution time fell from 11 minutes to under 2. Klarna put the 2024 profit improvement at $40M.
Why that qualifies and a copilot rollout does not: the assistant was the first responder. Chats landed on it, not on a person with a suggestion box. Humans became the escalation path.
Read the second-order numbers too. Repeat-inquiry rates were down 25% — a quality metric, not a volume one. A throughput number with no quality metric next to it is an unfinished claim, and it is usually unfinished on purpose.
And note where the boundary was drawn: customer service chat. Not credit decisions. Not the whole company.
How to tell whether something qualifies
Three tests. All three, or it is AI-enabled with good PR.
1. The default-assignment test. Open the routing config. If a work item goes to an agent first and to a human only on a named exception condition, it passes. If a human picks it up and calls a model, it fails. This is readable in code and queue rules, never in a strategy memo. Headcount is not the test — a 12-person company where 12 people do the work is not AI-first, and a 400-person company that routes tier-1 support to an agent is, for that function.
2. The meter test. You must be able to state cost per unit of work. Anthropic lists Claude Sonnet 4.5 at $3 per million input tokens and $15 per million output tokens; at those rates a 40-turn agent run averaging 15,000 input and 1,500 output tokens per turn costs about $2.70 in tokens alone, before retries. If you cannot say what a resolved ticket costs, you are not running the function. You are subsidising it. Instrumentation-first sequencing is the standing recommendation in agent engineering — build the meter before the agent, because that order matters more than the model choice. OpenTelemetry's GenAI semantic conventions give you the span and token-usage attributes to do this without inventing a schema.
3. The containment test. Unbounded loops are the most common production failure in agent systems, and the fix is a step cap enforced outside the model — the OpenAI Agents SDK takes a max_turns argument on Runner.run and raises MaxTurnsExceeded when the runner hits it. A cap the model is asked to respect in its prompt is not a cap. Same principle for spend, for writes to production data, for anything irreversible.
I am building a zero-human company and I still would not hand an agent unbounded spend authority today. Not because the reasoning is too weak. Because the blast radius of a loop with a credit card attached is the one failure you cannot roll back, and the choice of who picks the next step is exactly where that authority lives.
Sources
- Alphabet Founders' Letter 2016 (Sundar Pichai) — April 2016 letter in which Pichai describes the shift 'from mobile first to an AI first world' — the origin of the phrase in corporate use.
- Klarna AI assistant handles two-thirds of customer service chats in its first month — Primary source for the worked example: 2.3M conversations, two-thirds of chats, work equivalent of 700 agents, resolution time 11 min to under 2 min, ~$40M USD estimated profit impact for 2024.
- OpenAI Agents SDK — Running agents — Documents max_turns and the MaxTurnsExceeded exception: a step cap enforced by the runner, outside the model.
- OpenTelemetry Semantic Conventions for Generative AI — Standard span and metric attributes for agent runs, including token usage — the basis for per-unit cost measurement.
- Anthropic pricing — Published per-million-token input and output prices used to anchor cost-per-unit-of-work arithmetic.