Gil Allouche
← All writing
GlossaryRunning a company with AI

What an AI-first company is, and what it is not

The term's 2016 origin at Google, the two different things people mean by it, and three tests that separate AI-first from AI-enabled.

September 30, 2026·Gil Allouche·6 min read
A glossary entry. The definition sits at the top in one sentence, with the context that makes it useful underneath.

AI-first company

AI-first company
An AI-first company is one where AI systems are the default doer of routine work and humans handle the exceptions, not one where employees use AI tools to work faster.

Two words carry the load: default and doer. Default means the work routes to a model before it routes to a person — configuration, not intention. Doer means the system produces the output of record: the reply that ships, the invoice that posts, the campaign that goes live. Not a draft a human retypes.

Everything else in the term is decoration.

Where the term came from: Google, April 2016

Sundar Pichai used it in the Alphabet Founders' Letter published April 2016, describing a shift "from mobile first to an AI first world." The construction borrows from "mobile-first," a web design phrase in common use from roughly 2009: design for the constrained device first, then scale up, instead of shrinking a desktop layout.

The 2016 version was a product claim. Build the thing so the model is the interface. It said nothing about who inside the company did the work.

The operational meaning attached later, once tool-using agents could complete multi-step tasks without a human stitching the steps together. That drift is why the term needs disambiguation before it is usable in a sentence.

Two meanings. Label which one you mean

Sense 1: AI-first productSense 2: AI-first operating model
ClaimThe model is the core of what we sellThe model does the internal work
EvidenceModel in the product's critical pathWork items assigned to agents by default
Fails whenModel quality dropsNobody owns the exception queue
HeadcountNormalDeliberately flat
Typical era2016 onward2023 onward

Metadata.io was sense 1. The product ran on models, I hold 6 patents in AI-driven marketing, and we went from $0 to $15M ARR on $50M raised. The company that built it was a normal company: people did the work. Calling that "AI-first" in 2019 was accurate under sense 1 and would have been a lie under sense 2.

Someone searching this term today wants sense 2. Say which one you mean in the board deck and on the careers page — the two have almost no operational overlap.

What it is not: AI-enabled

AI-enabled means humans still own the work and use models to go faster. A support rep with a suggested-reply panel. A seller with a research copilot. An engineer with autocomplete. Output goes up, the assignment of work does not change, and if you switch the tool off on a Tuesday everything still ships — slower.

AI-first changes the assignment. Switch it off and the queue stops.

AI-enabled is not a rung on the ladder to AI-first. It is a different bet — buy tools, keep the org chart — and it is a defensible one. What is not defensible is making that bet and describing it as a transition.

AI-enabled routes work to a human who uses an AI tool; AI-first routes work to an agent and escalates only exceptions to a humanAI-enabledWork arrivesHuman does itAI assistsShippedAI-firstWork arrivesAgent does itChecks pass?ShippedHuman: exception

Two other neighbours get confused with it. AI-native is a founding condition, not an operating one — a company started after late 2022 whose first architecture assumed a model in the loop. A 40-year-old insurer can become AI-first. It cannot become AI-native. Digital transformation is about moving paper processes into software; AI-first assumes the software already exists and asks who operates it.

A worked example with real numbers: Klarna, 2024

Klarna published the cleanest documented case of sense 2. In February 2024 the company said its OpenAI-powered assistant had, in its first month live, handled 2.3 million conversations — about two-thirds of its customer service chats — doing the work equivalent of 700 full-time agents across 23 markets and 35 languages. Average resolution time fell from 11 minutes to under 2. Klarna put the 2024 profit improvement at $40M.

Why that qualifies and a copilot rollout does not: the assistant was the first responder. Chats landed on it, not on a person with a suggestion box. Humans became the escalation path.

Read the second-order numbers too. Repeat-inquiry rates were down 25% — a quality metric, not a volume one. A throughput number with no quality metric next to it is an unfinished claim, and it is usually unfinished on purpose.

And note where the boundary was drawn: customer service chat. Not credit decisions. Not the whole company.

How to tell whether something qualifies

Three tests. All three, or it is AI-enabled with good PR.

1. The default-assignment test. Open the routing config. If a work item goes to an agent first and to a human only on a named exception condition, it passes. If a human picks it up and calls a model, it fails. This is readable in code and queue rules, never in a strategy memo. Headcount is not the test — a 12-person company where 12 people do the work is not AI-first, and a 400-person company that routes tier-1 support to an agent is, for that function.

2. The meter test. You must be able to state cost per unit of work. Anthropic lists Claude Sonnet 4.5 at $3 per million input tokens and $15 per million output tokens; at those rates a 40-turn agent run averaging 15,000 input and 1,500 output tokens per turn costs about $2.70 in tokens alone, before retries. If you cannot say what a resolved ticket costs, you are not running the function. You are subsidising it. Instrumentation-first sequencing is the standing recommendation in agent engineering — build the meter before the agent, because that order matters more than the model choice. OpenTelemetry's GenAI semantic conventions give you the span and token-usage attributes to do this without inventing a schema.

3. The containment test. Unbounded loops are the most common production failure in agent systems, and the fix is a step cap enforced outside the model — the OpenAI Agents SDK takes a max_turns argument on Runner.run and raises MaxTurnsExceeded when the runner hits it. A cap the model is asked to respect in its prompt is not a cap. Same principle for spend, for writes to production data, for anything irreversible.

I am building a zero-human company and I still would not hand an agent unbounded spend authority today. Not because the reasoning is too weak. Because the blast radius of a loop with a credit card attached is the one failure you cannot roll back, and the choice of who picks the next step is exactly where that authority lives.

Sources

  1. Alphabet Founders' Letter 2016 (Sundar Pichai) — April 2016 letter in which Pichai describes the shift 'from mobile first to an AI first world' — the origin of the phrase in corporate use.
  2. Klarna AI assistant handles two-thirds of customer service chats in its first month — Primary source for the worked example: 2.3M conversations, two-thirds of chats, work equivalent of 700 agents, resolution time 11 min to under 2 min, ~$40M USD estimated profit impact for 2024.
  3. OpenAI Agents SDK — Running agents — Documents max_turns and the MaxTurnsExceeded exception: a step cap enforced by the runner, outside the model.
  4. OpenTelemetry Semantic Conventions for Generative AI — Standard span and metric attributes for agent runs, including token usage — the basis for per-unit cost measurement.
  5. Anthropic pricing — Published per-million-token input and output prices used to anchor cost-per-unit-of-work arithmetic.

Related