Gil Allouche
← All writing
GlossaryEntrepreneurship in the age of AI

What counts as an AI startup, and what does not

The model-removal test, the two senses the term carries, and the cost line that tells you whether a company is actually an AI startup.

October 3, 2026·Gil Allouche·10 min read
A glossary entry. The definition sits at the top in one sentence, with the context that makes it useful underneath.

AI startup

AI Startup
An AI startup is a new company whose core product only works because of machine-learning models — remove the models and there is nothing left to sell.

That test is subtractive, and it is the only one that survives contact with a pitch deck. Every other definition in circulation — "uses AI", "AI-powered", "built on LLMs" — is satisfied by a company that would keep billing customers if every model it calls went dark tomorrow. Those are marketing claims wearing a category label.

Where the term came from

The category predates ChatGPT by a decade. AlexNet won ImageNet in 2012 and made deep learning the default approach to perception problems; Google bought DeepMind in 2014. From roughly 2012 to 2020, "AI startup" meant a research team that trained its own model and sold it into one vertical: computer vision for inspection, speech for call centres, gradient-boosted fraud scoring. You needed a PhD on the founding team and a labelled dataset nobody else had.

The meaning moved in late 2022. Once a hosted general-purpose model was available on a credit card, a founder could ship an AI product without training anything — which collapsed the old barrier to entry and broke the old definition in the same stroke. OpenAI's GPT-3 API had been in limited beta since June 2020, but it was the ChatGPT launch on 30 November 2022, followed by GPT-4 in March 2023, that turned frontier-model access into an ordinary line item on a corporate card. The plumbing followed: Anthropic published the Model Context Protocol in November 2024, standardising how a model reaches tools, files and databases. Stanford's 2025 AI Index counted $252.3 billion in global private AI investment for 2024, of which $109.1 billion was US-based — roughly twelve times China's $9.3 billion. Both senses of the term below are chasing that money, and the pitch decks rarely say which one they mean.

The shift also changed who founds these companies. When the hard part was training, the founding team was selected for research ability and data access; when the hard part became making a rented model behave reliably inside someone else's workflow, the scarce skills moved to evaluation, product design and domain knowledge. The old definition quietly assumed the two were the same thing. They are not, and conflating them is why a 2015 taxonomy of the category reads as obsolete today.

Two senses, labelled

Sense 1 — AI-as-productSense 2 — AI-as-operator
What the customer buysModel output: a draft, a decision, a completed taskAn ordinary product, AI or not
Where the AI sitsIn the request path, billed as cost of goodsIn the company's internal workflows
Confused withAn AI feature on a pre-AI productBeing "lean" or "automated"
Fails whenModel quality or token cost movesThe agents need a human for the step that matters
Headcount signalNone. Can be 200 peopleSmall by construction

Both at once is rare. Sense 2 is what I am building publicly as a zero-human company, and it is a different claim from Sense 1 — I have written about the distinction in what an AI-first company is, and what it is not and AI-first means AI does the work by default.

Two senses of "AI startup" on one gridClassic SaaSneither senseAI is the productSense 1Agents run the companySense 2Both sensesrare, and claimed oftenModel sits in the product →Agents do the work →

What it is not: an AI-enabled company

Notion sells per-seat workspaces and bundles Notion AI into its paid plans. Cut the AI and Notion is still a database, a wiki and an editor that people pay for. That is a feature decision, not a company definition, and the neighbouring term for it is AI-enabled.

It is also not an AI lab. Anthropic, OpenAI and Mistral train frontier models and sell access to them — a capital-intensive research business with a cost structure that has almost nothing in common with an application company buying tokens by the million. Public estimates from Epoch AI put the compute bill for a single frontier training run in the tens to hundreds of millions of dollars; an application company's largest model-side commitment is usually a monthly invoice it can cancel. Treating the two as one category is how people end up comparing a $20 seat to a training run.

Being a "wrapper" disqualifies nothing. Calling someone else's model still passes the subtraction test; what matters is whether the prompts, the retrieval layer, the eval suite and the escalation rules are assets you own. Model endpoints are not permanent — providers deprecate versions on published schedules, and a product pinned to one model name inherits that calendar — so the durable asset is the harness that lets you swap the model and re-run the evals in an afternoon. If a customer could get the same result by pasting their question into a chat window, you do not have a product. You have a bookmark.

Edge cases the test has to survive

Three shapes make subtraction harder to apply than it sounds, and all three turn up in diligence.

The first is partial degradation. Switch off the model in a recruiting tool and you may still have an applicant tracking database: worse, but sellable. The question is not whether anything survives, it is whether what survives is what the customer agreed to pay for. If the contract prices shortlists delivered, the database is packaging.

The second is a model the customer never sees. Demand forecasting inside a logistics operator, or risk scoring inside a lender, can be the entire basis of the offer while having no interface at all. Subtraction asks about the deliverable, not about whether there is a chat box.

The third is the human in the loop. Plenty of products route model output past a reviewer before it ships. That does not disqualify them; it only matters where the economics sit. If the reviewer's time, not the model's output, is what scales with revenue, the company is a services business with a model attached.

Worked example, with the bill attached

Take a support-deflection agent sold per resolved ticket. Clearest Sense 1 shape there is, because the deliverable is model output and nothing else.

A conversation runs 40 turns. Each turn sends roughly 15,000 input tokens (system prompt, retrieved docs, conversation history) and returns 1,500. At Claude Sonnet list prices of $3 per million input tokens and $15 per million output tokens, that is $0.045 + $0.0225 = $0.0675 per turn, or $2.70 per conversation. Five thousand conversations a month is $13,500 in inference before you have paid for anything else. Smaller models change the answer by an order of magnitude — GPT-4o mini lists at $0.15 per million input and $0.60 per million output tokens, a twentieth of Sonnet's input price — which makes model selection a pricing decision, not an engineering preference.

What would this cost you?
40
15,000
5,000
per run
$2.70
full pass
$13,500

Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.

Prompt caching moves the same number without changing models. If 12,000 of those 15,000 input tokens are a stable system prompt plus retrieved docs, and cache reads bill at a tenth of the base input rate, the per-turn cost becomes $0.0036 + $0.009 + $0.0225 = $0.0351 — $1.40 per conversation, or $7,020 a month at the same volume. That is the difference between a viable and an unviable price per resolved ticket: sell this agent at $1 a ticket and the uncached version loses $1.70 on every one.

Cursor sits at the other end of the same category: a Pro seat is $20 per month, and the completions a heavy user consumes are the product. Different shape, identical structural fact. Inference sits on the gross margin line and moves with usage. If your model spend lives in the R&D budget rather than COGS, you are AI-enabled and your margins are about to tell you so. I wrote the build-order version of this in build a SaaS app with AI, starting with the token meter and how to make an AI startup, starting with the meter.

What a real COGS line changes

Software businesses were built on a marginal cost near zero; public SaaS benchmarks have long put gross margins in the 70–80% band for exactly that reason. Inference breaks the assumption. Every successful use of the product costs money, so under a flat seat price the most engaged customers are the least profitable ones. Three consequences follow, and they are commercial rather than technical:

  • Pricing has to track the unit that correlates with tokens — resolved tickets, generated documents, accepted completions — or a seat price needs a usage ceiling behind it.
  • Capacity becomes a customer-facing concern. Provider rate limits, latency spikes and deprecation windows land on the customer's workflow, so fallback routing is part of the product, not the infrastructure backlog.
  • Margin is earned by engineering as well as by scale. Caching, context trimming, routing easy turns to cheap models and retrying only when an eval fails are the levers that move the number above. That is why the cached-versus-uncached arithmetic is not a footnote.

Defensibility when the model is rented

If the model is a commodity on tap, the durable assets are everything around it: a proprietary eval set that encodes what "good" means in this domain, a failure taxonomy built from real incidents, permissioned access to data the customer will not hand to a general-purpose chatbot, and integration depth that makes the output usable without copy-paste. None of that is novel strategy — it is switching cost and distribution, as before. The part specific to this category is that a harness making model swaps cheap turns provider progress from a threat into a tailwind: each new model release is a margin or quality upgrade you can ship in a day instead of a migration you have to budget for.

How to tell whether something qualifies

Four questions. The first is definitional; the rest tell you whether the thing is real.

  1. Subtraction. Switch off every model call. Is there a sellable product left? If yes, it is AI-enabled. If no, it is an AI startup.
  2. Where the cost sits. Find token spend on the P&L. COGS means the model is in the product. Opex means it is in the office.
  3. Evals before opinions. A real AI product team can tell you its pass rate on a fixed test set and what happened the last time it changed models. No eval suite means nobody knows whether the product got worse this week.
  4. Control outside the model. Unbounded loops and runaway tool calls are the most common production failure in agent systems, and the fix is never a better prompt — it is a step cap, a spend cap and a tool allow-list enforced in code the model cannot edit. Anthropic's December 2024 guidance on building effective agents makes the same point, separating fixed workflows from open-ended agent loops and recommending the simplest structure that solves the task. A widely documented pattern is to split planning, execution and review across separate agents so no single loop both decides and acts unchecked, with the caps held outside every one of them: build a three-agent system that stops when it should.

For Sense 2 the test is narrower and much harder to fake. Name the function that no human touches end to end, and say who holds spend authority. I would not hand an agent unsupervised spend authority yet, and a company claiming to be agent-run while a person still approves every outbound message is describing a tool, not an operating model.

Sources

  1. Stanford HAI, 2025 AI Index Report — Global private AI investment figure for 2024
  2. Anthropic — Introducing the Model Context Protocol — November 2024 publication of MCP as the tool/data standard
  3. Anthropic — Model pricing — Claude Sonnet input/output token prices used in the worked example
  4. OpenAI — API pricing — Cheaper per-token tiers referenced against the Sonnet math
  5. Anthropic Engineering — Building effective agents — Documented agent loop failure modes and the case for external control
  6. Cursor pricing — Per-seat price for an AI-as-product company where inference is COGS
  7. Notion pricing — Example of AI bundled into a product that predates it
  8. ImageNet Classification with Deep Convolutional Neural Networks (AlexNet) — Dates the deep-learning turn that created the venture category in 2012

Related