Gil Allouche
← All writing
ReferenceRunning a company with AI

Starting a business with AI agents, and what stays human

Where agents earn their keep in a new company, what a run actually costs, the failure modes to design for before launch, and the three things I would not delegate.

October 7, 2026·Gil Allouche·9 min read
A reference explainer. Every factual claim links to its source; where this is an opinion from operating experience, it says so.

Anthropic's prompt-caching docs price a cache read at 0.1× the base input token rate and a cache write at 1.25×. OpenAI's Batch API takes 50% off standard rates if you can wait up to 24 hours for results. Token cost is a solved engineering problem. It is not the thing that decides whether your company works.

Starting a business with AI agents means writing your process down as code, giving agents only the tasks where a wrong answer is cheap and reversible, and keeping three things outside the model: spend authority, irreversible writes, and anything a customer sees at volume without review.

I took Metadata.io from $0 to $15M ARR with a human company — $50M raised, six patents in AI-driven marketing. I am building the next one as a zero-human company, with agents doing the work and me paying the inference bills. The constraint has moved. It is no longer "can an agent do this task." It is "what happens on the 400th run, the one where it does the task wrong and nobody is watching."

The jobs agents hold down in month one

The pattern in every job that works: there is an oracle. Something outside the model can say that output was wrong without a human reading it.

JobWhy it worksWhat breaks first
Inbound lead research and enrichmentBounded input, checkable output, a wrong note costs almost nothingSource APIs throttle — GitHub's REST API caps authenticated callers at 5,000 requests/hour, 60/hour unauthenticated
Support triage and draft repliesHistorical tickets are labelled ground truthInvented policy stated as fact to a customer
Code changes behind a test suiteThe tests are the oracleThe agent edits the test instead of the code
Data reconciliation and cleanupDiffs are verifiable, runs can be replayedNon-idempotent writes creating duplicates on retry
Content and proposal drafts to a named editorA human gates publicationVolume with no reviewer capacity behind it

Jobs with no oracle — pricing decisions, hiring, anything where "plausible" and "correct" look identical on the screen — are where agent projects quietly rot. The output reads fine for six weeks. Then you find out the segmentation has been wrong since day one.

Start with the job where you already know what good looks like. Replacing a sales team with AI is this argument applied to one function.

The architecture I would build, and the one I would not

I would not build a role-play company: a CEO agent that delegates to a CMO agent that delegates to an SDR agent. It demos well. In production, every handoff is a place to lose state, cost scales linearly with the number of agents in the conversation, and accuracy does not improve to match. You end up debugging a group chat.

Build the opposite shape. Plain code holds the process. Agents are called for the steps that need judgement, one narrow job each, and the code holds the caps.

Where the caps live in an agent loopTriggerOrchestratorStep capSpend capIdempotency keyPlain codeAgent turnTool callWrite gateExecuteHuman queuecountedreversibleirreversible

Every turn is counted by the orchestrator, not by the model. A model asked to limit itself will cheerfully report that it is limiting itself while looping. The frameworks already give you the primitive: the OpenAI Agents SDK takes a max_turns argument on Runner.run and raises MaxTurnsExceeded; LangGraph ships a default recursion_limit of 25 and raises GraphRecursionError past it. Use them. If you are hand-rolling, the counter is five lines, and it belongs in your code, not your prompt.

Tools reach the agent through a declared interface. Model Context Protocol, published by Anthropic in November 2024, is the version of that interface you do not have to invent. It gives you one place to see what an agent can touch — which is also the list you audit when something goes wrong.

I would not reach for a framework before writing one loop by hand. You need to have felt the retry logic to know what the framework is hiding. What an AI agent framework actually is covers what is and is not in the box, and build a three-agent system that stops when it should is the shape above with the stopping conditions written out.

What a run costs, and what the bill actually is

Take a concrete agent: inbound lead research. Twelve tool-using turns per account — fetch the site, read two pages, check the CRM, write a summary. Roughly 15,000 input tokens per turn once you count the system prompt, prior messages and page text. About 900 output tokens per turn. Three thousand accounts a month.

What would this cost you?
40
15,000
5,000
per run
$2.70
full pass
$13,500

Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.

That figure is the uncached worst case. Prompt caching cuts most of it, because the system prompt and tool schemas are identical on every turn: Anthropic prices cache reads at 0.1× base input, and anything you can run asynchronously qualifies for OpenAI's 50% Batch discount with a 24-hour window.

Now the part the calculator does not show.

Stripe takes 2.9% + 30¢ on a successful domestic card charge. A $19 product gives up about 85¢ per transaction before anyone looks at a token bill. Your observability vendor, your error tracker and your database have monthly floors that do not care how few runs you did. And if a human reviews output, their hour is the dominant line item at any volume that matters — 3,000 runs at 20 seconds of review each is over 16 hours of someone's month.

So the number to carry around is not cost per thousand tokens. It is all-in cost per completed outcome, including review, set against what someone pays for that outcome. If you want the meter wired in from the first commit rather than bolted on at month three, build a SaaS app with AI starting with the token meter.

Five failure modes to design for before launch

None of these are exotic. All of them will show up.

Unbounded loops. The agent calls a tool, dislikes the result, calls it again, forever. This is the most common way an agent run turns into a four-figure bill. The fix is a step cap enforced outside the model and a hard wall-clock timeout on the whole run.

Non-idempotent writes. A timeout is not a failure; it is an unknown. The agent retries, and you get two charges, two emails, two records. Stripe documents idempotency keys for exactly this, and the pattern generalises: every write an agent can make should carry a caller-generated key that makes a replay a no-op.

Prompt injection through fetched content. OWASP ranks prompt injection as the first risk class in its Top 10 for LLM Applications. Any agent that reads a web page, a PDF or an inbound email is reading attacker-controlled text. Treat fetched content as data, never as instructions, and give the agent no tool whose misuse you could not survive.

Rate limits and partial failure. Documented ceilings exist everywhere — GitHub's 5,000 requests/hour for authenticated callers is typical. An agent that does not handle 429 with backoff will burn its step budget on retries and return a confident, half-sourced answer. "Hit a rate limit and stopped" belongs in the design as a first-class, loggable terminal state, distinct from both success and an unexplained crash.

Silent drift. No oracle, no complaints, no signal. The only defence is a frozen set of 30–50 real inputs with known-good outputs that you re-run on every prompt or model change. Not a vibe check. A file in the repo.

What stays human, and where the law agrees

Spend authority. I would not give an agent an unconstrained payment method. If you do delegate spend, delegate it through something that cannot overspend: Stripe Issuing supports hard spending limits by interval and allowed merchant categories, so the ceiling is enforced by the card network rather than by your prompt.

Decisions about people. GDPR Article 22 gives individuals a right not to be subject to decisions based solely on automated processing where those decisions have legal or similarly significant effects. Hiring, credit, access, termination. An agent can prepare the file. A person decides and signs.

Disclosure. EU AI Act Article 50 requires that people be informed when they are interacting with an AI system, unless it is obvious. Build it into the first message, not into a footer nobody reads.

Volume outbound. I would not point an agent at cold email at scale. Google's sender guidelines require authentication for bulk senders and tell you to keep reported spam below 0.3% — a rate an eager agent can blow through in a single afternoon, taking your sending domain with it. Agents are good at researching the account and drafting the thing. The send decision stays gated. That split is the whole argument in what to automate in AI account-based marketing, and what to gate.

The parts no agent touches

Incorporation, bank KYC, signing a contract, taking the liability. An agent can fill a form. It cannot be the person whose identity the bank verifies or whose signature binds the company. Plan on doing that work yourself, and plan on it taking longer than the software.

You are not building an autonomous business. You are building a business where one human holds the irreversible decisions and agents hold the repeatable work, and where the boundary between those two is written down in code you can read. What counts as an AI startup, and what does not is where I draw that line in public.

Here is the number that decides it. Per-token prices keep falling and caching keeps cutting them further — the compute side of your cost curve is going the right way without you doing anything. Your cost to acquire a customer is not. If an agent can produce a completed outcome for cents and you still cannot get it in front of someone who will pay for it, you do not have a cost problem. You have the same distribution problem every company has had, and no agent is going to hand it to you solved.

Sources

  1. Anthropic — Prompt caching — Cache reads priced at 0.1x base input tokens, cache writes at 1.25x
  2. OpenAI — Batch API guide — 50% discount vs synchronous pricing, results within a 24-hour window
  3. OpenAI Agents SDK (Python) repository — max_turns argument on Runner.run and the MaxTurnsExceeded exception
  4. LangGraph documentation — Default recursion_limit of 25 and GraphRecursionError when a graph exceeds it
  5. Model Context Protocol specification — Open protocol for exposing tools and data to models, introduced by Anthropic in November 2024
  6. GitHub REST API rate limits — 5,000 requests/hour authenticated, 60/hour unauthenticated — a documented ceiling agents hit
  7. OWASP Top 10 for LLM Applications — Prompt injection listed as the first risk class for LLM applications
  8. Stripe — API idempotent requests — Idempotency keys as the documented mechanism for safe retries
  9. Stripe — Issuing spending controls — Hard spending limits by interval and merchant category on issued cards
  10. Stripe pricing — 2.9% + 30¢ per successful domestic card charge in the US
  11. Google — Email sender guidelines — Bulk sender authentication requirements and the sub-0.3% spam-rate threshold
  12. GDPR Article 22 — Right not to be subject to solely automated decisions with legal or similarly significant effects
  13. EU AI Act Article 50 — Transparency obligation to inform people they are interacting with an AI system

Related