Start a company with AI: one agent, three hard caps
Build a runnable Python agent that triages inbound leads, halts at a step cap and a spend ceiling, logs every token, and cannot send anything without you.
What you'll have, and what you need
operator.py is one Python file, about 120 lines, and it stops dead at either 8 model calls or $0.25 of spend. It reads a queue of inbound signups, decides which ones match your ICP, and writes a decision plus a draft reply to disk. Budget 30 minutes if you already have Python and an OpenAI key. Most of that half hour is pasting one file and then deliberately breaking it to watch the cap fire.
You need Python 3.10 or newer and a terminal. You need an OpenAI API key from platform.openai.com/api-keys, on an account with credit — the script defaults to gpt-4.1-mini, listed at $0.40 per 1M input tokens and $1.60 per 1M output tokens on the pricing page as of writing. Check it; prices move. And you need your ICP written down in one paragraph. Not a persona deck. A paragraph you could read aloud. No framework. One pip install openai.
Day one is a process, not a stack
I took Metadata.io from $0 to $15M ARR with humans doing most of the operating work, raised $50M against it, and came out with 6 patents in AI-driven marketing. Before that, Silver Spotfire, $50K to $1.5M ARR. I am building the next one as a zero-human company — AI agents do the work, and I pay the bills for them.
The first artifact is not a model choice, and it is not an agent framework. It is one business process, written down tightly enough that a loop can execute it and you can audit what it did. Pick the framework first and you inherit someone else's control flow, which is precisely where the step cap and the spend ledger have to live.
Lead triage is the right first process: high volume, reversible, and a wrong answer costs you an email instead of a contract. So that is what we are building. The longer argument about what "AI-first" does and doesn't mean is here.
The whole file: operator.py
Two files. First, the queue — note the reserved .example domains, and use those while testing so you never email a real person by accident:
[
{"email": "rita@acme.example", "company": "Acme Analytics", "domain": "acme.example", "employees": 140, "note": "wants to replace 3 SDRs with agents"},
{"email": "ops@northwind.example", "company": "Northwind", "domain": "northwind.example","employees": 12, "note": "asked about pricing tiers"},
{"email": "student@school.test", "company": "n/a", "domain": "school.test", "employees": 1, "note": "class project, needs free access"}
]
Then the operator itself:
#!/usr/bin/env python3
"""operator.py - one agent, three hard caps, safe to put on cron."""
import json, os, sys, time, pathlib
from openai import OpenAI
MODEL = "gpt-4.1-mini"
PRICE_IN, PRICE_OUT = 0.40, 1.60 # USD per 1M tokens - verify on the pricing page
MAX_STEPS = 8 # model calls per lead, enforced in Python
MAX_USD_PER_RUN = 0.25 # spend ceiling for the whole invocation
OUT = pathlib.Path("out"); OUT.mkdir(exist_ok=True)
AUDIT, APPROVALS = OUT / "audit.jsonl", OUT / "approvals.jsonl"
client = OpenAI() # reads OPENAI_API_KEY from the environment
spend = 0.0
ICP = ("We sell to B2B SaaS companies with 50-500 employees that run paid demand "
"generation and have at least one marketing ops person. Not agencies, not "
"students, not companies under 20 people.")
SYSTEM = f"""You triage inbound signups. ICP: {ICP}
Procedure, in order:
1. Call lookup_crm with the domain.
2. If you would send anything outbound, call request_human_approval. You have no send tool.
3. Call save_decision exactly once with verdict in {{qualified, nurture, reject}}.
Then reply with one sentence and stop. Do not repeat a tool call with identical arguments."""
CRM = {"northwind.example": {"status": "existing customer", "arr_usd": 24000}}
def audit(event, **fields):
rec = {"ts": round(time.time(), 3), "event": event, **fields}
with AUDIT.open("a") as f:
f.write(json.dumps(rec) + "\n")
print(f"[{event}] " + " ".join(f"{k}={v}" for k, v in fields.items()))
def lookup_crm(domain: str):
return CRM.get(domain, {"status": "not in crm"})
def save_decision(domain: str, verdict: str, reason: str, draft_reply: str):
path = OUT / f"{domain}.json"
path.write_text(json.dumps(
{"domain": domain, "verdict": verdict, "reason": reason,
"draft_reply": draft_reply, "model": MODEL}, indent=2))
return {"saved": str(path)}
def request_human_approval(action: str, detail: str):
with APPROVALS.open("a") as f:
f.write(json.dumps({"ts": round(time.time(), 3), "action": action,
"detail": detail, "status": "PENDING"}) + "\n")
return {"status": "PENDING_APPROVAL", "executed": False,
"note": "A human must approve this. Continue without it."}
def _tool(name, desc, props):
return {"type": "function", "function": {
"name": name, "description": desc,
"parameters": {"type": "object", "properties": props,
"required": list(props), "additionalProperties": False}}}
S = {"type": "string"}
TOOLS = [
_tool("lookup_crm", "Look up a domain in the CRM.", {"domain": S}),
_tool("request_human_approval", "Queue an action for human approval. Nothing is sent.",
{"action": S, "detail": S}),
_tool("save_decision", "Persist the triage decision. Call once.",
{"domain": S, "verdict": S, "reason": S, "draft_reply": S}),
]
DISPATCH = {"lookup_crm": lookup_crm, "save_decision": save_decision,
"request_human_approval": request_human_approval}
def usd(usage):
return usage.prompt_tokens / 1e6 * PRICE_IN + usage.completion_tokens / 1e6 * PRICE_OUT
def run_lead(lead):
global spend
messages = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": json.dumps(lead)}]
for step in range(1, MAX_STEPS + 1):
if spend >= MAX_USD_PER_RUN:
audit("spend_cap_hit", spend_usd=round(spend, 4))
return "HALTED: spend cap"
r = client.chat.completions.create(model=MODEL, messages=messages, tools=TOOLS)
spend += usd(r.usage)
audit("model_call", step=step, in_tok=r.usage.prompt_tokens,
out_tok=r.usage.completion_tokens, spend_usd=round(spend, 4))
msg = r.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
return msg.content or "(no content)"
for tc in msg.tool_calls:
name = tc.function.name
try:
args = json.loads(tc.function.arguments or "{}")
except json.JSONDecodeError as e:
args, result = {}, {"error": f"unparseable arguments: {e}"}
else:
fn = DISPATCH.get(name)
result = fn(**args) if fn else {"error": f"unknown tool: {name}"}
audit("tool_call", step=step, tool=name, result=json.dumps(result)[:120])
messages.append({"role": "tool", "tool_call_id": tc.id,
"content": json.dumps(result)})
audit("step_cap_hit", steps=MAX_STEPS)
return "HALTED: step cap"
if __name__ == "__main__":
path = sys.argv[1] if len(sys.argv) > 1 else "inbound.json"
leads = json.loads(pathlib.Path(path).read_text())
for lead in leads:
audit("lead_start", email=lead["email"])
print(" ->", run_lead(lead))
audit("run_complete", leads=len(leads), spend_usd=round(spend, 6))
There is no send_email function in that file. That is the design, not an omission. An agent cannot do what it has no tool for, and a missing tool beats an instruction in a prompt every time.
Build it
- Make the project and a virtual environment:
mkdir operator && cd operator, thenpython3 -m venv .venv, thensource .venv/bin/activate(Windows:.venv\Scripts\activate). Your prompt should now be prefixed with(.venv). - Install the SDK:
pip install openai. pip ends with a line beginningSuccessfully installedthat includesopenai-1.and its dependencies —httpx,pydantic,anyio. - Export your key:
export OPENAI_API_KEY="sk-..."from platform.openai.com/api-keys. Verify withpython -c "import os; print(bool(os.getenv('OPENAI_API_KEY')))"— you wantTrue, and you do not want the key itself echoed into your shell history file. - Save the two files from the section above as
inbound.jsonandoperator.pyin that directory. Replace theICPstring with your paragraph. Leave the.exampleand.testaddresses alone for now. - Run it:
python operator.py. You should see interleaved[model_call]and[tool_call]lines, one->sentence per lead, and a final[run_complete]withspend_usd. - Read what it wrote:
cat out/acme.example.jsonfor the decision and the draft, andcat out/approvals.jsonlfor anything it wanted to send. Nothing inapprovals.jsonlhas been executed. - Now break it on purpose. Change
MAX_STEPS = 8toMAX_STEPS = 1, deleteout/audit.jsonl, and re-run. Every lead should end in[step_cap_hit] steps=1andHALTED: step cap. That is the proof the cap lives in Python and not in the prompt's good intentions. - Put it back to 8. Count your model calls per lead with
grep -c model_call out/audit.jsonl. That number, times your per-lead cost, is your unit economics for this process.
What you should see
This is the shape of a healthy run. Token counts and the order of tool calls will differ for your ICP and your leads:
[lead_start] email=rita@acme.example
[model_call] step=1 in_tok=... out_tok=... spend_usd=0.0004
[tool_call] step=1 tool=lookup_crm result={"status": "not in crm"}
[model_call] step=2 in_tok=... out_tok=... spend_usd=0.0009
[tool_call] step=2 tool=save_decision result={"saved": "out/acme.example.json"}
[model_call] step=3 in_tok=... out_tok=... spend_usd=0.0013
-> Acme Analytics is qualified: 140 employees, paid demand gen, agent replacement use case.
[lead_start] email=student@school.test
...
[run_complete] leads=3 spend_usd=0.004
out/audit.jsonl is one JSON object per line, append-only, and it is the only record of what happened. Three leads through this loop produce 8–12 lines in it. If a decision looks wrong a week later, the audit line tells you which tool returned what, and the decision file tells you what the model concluded from it.
One thing to note before you schedule this: MAX_USD_PER_RUN is a per-invocation ceiling, so the number that matters on cron is the ceiling multiplied by the number of invocations. Every fifteen minutes is 96 runs, or $24 a day at the cap. Either raise the cap and run less often, or persist spend to a file next to the audit log and read it back at startup so the ledger spans the whole day.
The four controls and where each one lives
| Control | Where it is enforced | Failure it prevents |
|---|---|---|
Step cap, MAX_STEPS = 8 | the for loop, before each model call | unbounded tool loops — the most common way an agent burns money without producing output |
Spend ceiling, MAX_USD_PER_RUN | usage.prompt_tokens / usage.completion_tokens summed after every call | a long context silently multiplying per-call cost across a batch |
| Approval gate | there is no send tool; request_human_approval writes PENDING and returns executed: False | the agent contacting a customer, or a student, on its own authority |
| Audit log | audit() on every model call and every tool call | having no way to reconstruct a bad decision |
Capping iterations outside the model is standard practice, not a trick. OpenAI's Agents SDK ships a max_turns parameter that raises when exceeded, and Anthropic's guidance on agent loops is to give them explicit stopping conditions and keep the composition simple. A prompt that says "don't loop" is a request. A range(1, MAX_STEPS + 1) is a fact. Stopping behaviour gets harder once agents can delegate to each other, because one agent's budget is spent by calls it does not make itself — the same cap-in-code principle then has to hold at the orchestration layer, which is the case worked through in this build.
When it doesn't work
No key, or the wrong key. Starting the script without the environment variable raises before any network call:
openai.OpenAIError: The api_key client option must be set either by passing
api_key to the client or by setting the OPENAI_API_KEY environment variable
A key that exists but is revoked, or scoped to a different project, returns a 401 instead: openai.AuthenticationError: Error code: 401 - {'error': {'message': 'Incorrect API key provided: sk-...'}}. Re-export the key in the same shell you run from — a new terminal tab does not inherit it — and confirm the key belongs to the project you funded.
No credit on the account. This one looks like rate limiting and is not:
openai.RateLimitError: Error code: 429 - {'error': {'message': 'You exceeded your
current quota, please check your plan and billing details.', 'type': 'insufficient_quota'}}
insufficient_quota means billing, not throughput. Add credit in the dashboard. A genuine rate limit has type: 'requests' or 'tokens' and is worth retrying; insufficient_quota never succeeds on retry, so never wrap it in a retry loop — you will pay for the latency and get the same error.
A 400 on the tool reply. Edit the loop and drop the messages.append(msg.model_dump(...)) line, or answer two tool calls with one tool message, and the next call fails:
openai.BadRequestError: Error code: 400 - {'error': {'message': "An assistant message with
'tool_calls' must be followed by tool messages responding to each 'tool_call_id'..."}}
Append the assistant message first, then exactly one {"role": "tool", "tool_call_id": ...} per entry in msg.tool_calls, including the ones your dispatcher rejected. An error string in the tool result is a valid response and lets the model recover. Silence is not.
Arguments that are not valid JSON. Tool arguments arrive as a string and are generated, so they can be truncated or malformed. The try/except json.JSONDecodeError branch exists for that: it turns the parse failure into a tool result the model can read and retry on the next step, and crucially it still appends a tool message, which keeps the conversation valid. Delete the handler and a single bad argument string raises mid-batch, abandoning the leads behind it. Because the audit log is append-only, a crashed run tells you exactly which lead to restart from.
Bad model id. openai.NotFoundError: Error code: 404 - {'error': {'message': 'The model ... does not exist or you do not have access to it'}}. Model names and access change. Take the exact string from the pricing page, not from memory.
It halts at the step cap every time. The cap did its job; the procedure is the problem. Look in out/audit.jsonl for the repeated tool_call line. An identical call twice means the tool returned something the model could not use — {"status": "not in crm"} reads as a failure unless you tell it that is a normal answer. Fix the return value, or fix the procedure. Raising MAX_STEPS to 20 converts a visible stall into an expensive one.
What I would not hand this agent yet
Spend authority. This loop can draft a reply and queue it; it cannot send, and it cannot touch a budget. That stays true until the approvals file has enough entries to score the agent's judgement against my own — read the queued actions, count how many you would have approved unchanged, and then automate only the class of action you approved every single time. Not the category. The class.
Second, write access to the CRM. save_decision writes a file in out/, which is reversible with rm. An agent with a CRM token writes to a system other humans and other automations read, and a bad field propagates into sequences, scoring and reporting before anyone notices. The full map of what I would automate versus gate in a go-to-market motion is here.
Third, the model picker. Leave gpt-4.1-mini in place until the process is correct, because changing models while the procedure is still wrong gives you two variables and no answer. Get the audit log boring first. Then swap the model and compare the same 20 leads: same decisions and lower spend_usd, keep it; different decisions, and you have learned the task is harder than the cheap model can handle, plus the price of that.
Your next unit is whatever process you are personally doing most often this week. Write it as a procedure, give it the smallest set of tools that can finish it, cap the loop, log every call. Then refuse to add a tool that does something you cannot undo with rm.
Sources
- OpenAI API pricing — Per-million-token input and output prices used in the cost ledger, including gpt-4.1-mini.
- OpenAI function calling guide — Tool schema format, the assistant/tool message sequence, and the tool_call_id requirement.
- OpenAI API keys — Where the reader gets OPENAI_API_KEY.
- openai-python (official SDK repository) — Install target, 1.x client, and the OPENAI_API_KEY environment variable default.
- OpenAI API error codes — 401 invalid authentication and 429 insufficient_quota error semantics quoted in the troubleshooting section.
- OpenAI Agents SDK documentation — max_turns as a documented, framework-level run cap — precedent for enforcing a step cap outside the model.
- Anthropic — Building effective agents — Agent loops need explicit stopping conditions and guardrails; favour the simplest composition that works.
- Python venv documentation — Virtual environment creation and activation commands.