Build a three-agent system that stops when it should
A runnable triage → enrichment → reply handoff chain in the OpenAI Agents SDK, with a step cap, a token ledger, and the three errors you will hit first.
What you will have running
Three agents, two function tools, one file. A triage agent routes an inbound email, an enrichment agent looks the sender up and names a pricing tier, a drafting agent writes the reply — wired as a directed handoff chain with a hard 8-turn cap and a token ledger printed at the end of every run. Budget 20 to 30 minutes, most of it reading the code rather than typing it.
You need Python 3.10 or newer, a terminal, and an OpenAI API key with billing enabled — create one at platform.openai.com/api-keys. The run is four or five calls to gpt-4.1-mini, which launched on 14 April 2025 at $0.40 per million input tokens and $1.60 per million output tokens, so about $0.01 of credit; check current prices yourself rather than trusting a number in a blog post. No vector database, no queue, no Docker.
Pick the topology before you write code
Two shapes cover almost everything people mean by "multi-agent system," and the OpenAI Agents SDK — shipped on 11 March 2025 as the production-ready successor to the experimental Swarm library — supports both. Choose first. They fail differently.
| Handoffs (directed graph) | Supervisor + agents-as-tools | |
|---|---|---|
| Who picks the next step | the agent currently running, by calling a handoff tool | one orchestrator, on every step |
| Conversation state | transfers to the receiving agent | stays with the orchestrator; the sub-agent is a function call that returns a string |
| SDK surface | Agent(handoffs=[other_agent]) | other_agent.as_tool(tool_name=..., tool_description=...) |
| Characteristic failure | two agents handing control back and forth | orchestrator context grows with every sub-agent result |
Handoffs are exposed to the model as tools named like transfer_to_<agent_name>, which is why the handoff_description field matters — it is the tool description the model reads when deciding. The as_tool() route keeps one agent in charge the whole time and pays for it in context: every sub-agent result lands back in the orchestrator's history and stays there for the rest of the run. Handoffs have the opposite knob — handoff(agent, input_filter=handoff_filters.remove_all_tools) from agents.extensions strips earlier tool calls out of the history the receiving agent inherits.
This tutorial builds the handoff version because the failure is easier to see. Who picks the next step is the whole definition of autonomy, and I wrote about that distinction separately in autonomous agents, defined by who picks the next step.
Anthropic reported that their multi-agent research system used roughly 15 times more tokens than chat interactions — and about 4 times more for single agents — while a Claude Opus 4 lead delegating to Claude Sonnet 4 subagents beat single-agent Claude Opus 4 by 90.2% on their internal research eval. Cognition published the opposite recommendation entirely — Walden Yan's "Don't Build Multi-Agents," June 2025 — on the grounds that splitting context across agents produces decisions that do not add up. Both are right. More agents buys you specialisation and costs you coherence and tokens, so three is where this chain stops.
Build it
- Create a virtualenv and install the SDK:
python -m venv .venv && source .venv/bin/activatethenpip install openai-agents. You should see a line endingSuccessfully installed openai-agents-…. Confirm withpip show openai-agents— it printsName: openai-agentsand aVersion:line; record that version somewhere, because the handoff and tracing APIs below are the ones you just installed, not the ones in the next release. - Set your key:
export OPENAI_API_KEY=sk-…, taken from platform.openai.com/api-keys. Verify without printing the secret:python -c "import os; print(bool(os.environ.get('OPENAI_API_KEY')))"should printTrue. If it printsFalse, you exported it in a different shell. - Save the file in the next section as
triage.py. Nothing to see yet — this is threeAgentobjects, two@function_toolfunctions, and amain()that prints the path the run took. - Run it:
python triage.py. You should see a--- path ---block containing twohandoff_call_itemlines, at least onetool_call_itemand its matchingtool_call_output_item, then a four-sentence reply, thenlast agent: Reply Drafterand amodel calls:line with token counts. - Force the cap:
python triage.py 2. The run stops before the drafter gets control and printsSTOPPED by cap:followed by aMaxTurnsExceededmessage naming the limit. This is the single most important line in the file — see the next section. - Open the trace. Tracing is on by default in the Agents SDK; go to platform.openai.com/traces and open the newest trace. You should see nested spans — one
agentspan per agent, ahandoffspan for each transfer,functionspans forlookup_accountandpricing_tierunder Enrichment, and agenerationspan per model call. Compare the span tree to your--- path ---output — they should tell the same story. SetOPENAI_AGENTS_DISABLE_TRACING=1if this run must send nothing. - Delete one agent. Replace the first line of
main()with a Python conditional that picks the starting agent from a keyword match, as shown after the code. Re-run. Themodel calls:count drops by one and onehandoff_call_itemdisappears from the path, because routing on three keywords never needed a language model.
triage.py
Complete and runnable. The sample data uses reserved documentation domains on purpose.
# triage.py - three agents, one directed handoff chain, one hard cap.
import asyncio
import json
import sys
from agents import Agent, Runner, function_tool
from agents.exceptions import MaxTurnsExceeded
# Stand-in for your CRM.
ACCOUNTS = {
"northwind.example": {
"company": "Northwind Retail",
"employees": 85,
"plan": "trial",
"owner": "ae-3@sales.example.com",
},
"acme.example.com": {
"company": "Acme Industrial",
"employees": 4200,
"plan": "enterprise",
"owner": "ae-1@sales.example.com",
},
}
@function_tool
def lookup_account(domain: str) -> str:
"""Look up an account by the sender's email domain.
Args:
domain: the part of the address after the @, e.g. northwind.example
"""
return json.dumps(ACCOUNTS.get(domain, {"found": False}))
@function_tool
def pricing_tier(employees: int) -> str:
"""Return the pricing tier for a company of this headcount.
Args:
employees: total headcount
"""
if employees < 50:
return "starter"
if employees < 500:
return "growth"
return "enterprise"
reply_drafter = Agent(
name="Reply Drafter",
handoff_description="Writes the final reply once the account facts are known.",
instructions=(
"You write replies to inbound sales email. Four sentences maximum. "
"Use only facts present in the conversation. If no account was found, "
"say so and ask one qualifying question. Never invent a price. "
"You are the last agent: produce the reply text and stop."
),
model="gpt-4.1-mini",
)
enrichment = Agent(
name="Enrichment",
handoff_description="Looks the sender up in the CRM and names the pricing tier.",
instructions=(
"Extract the sender's email domain. Call lookup_account once with it. "
"If an account comes back, call pricing_tier once with its employee count. "
"State what you found in one line, then hand off to Reply Drafter. "
"Do not write the reply yourself. Do not call either tool twice."
),
tools=[lookup_account, pricing_tier],
handoffs=[reply_drafter],
model="gpt-4.1-mini",
)
triage = Agent(
name="Triage",
instructions=(
"You route inbound email. If the message mentions pricing, seats, plans "
"or a renewal, hand off to Enrichment. Otherwise hand off to Reply Drafter. "
"Hand off immediately. Do not answer the email yourself."
),
handoffs=[enrichment, reply_drafter],
model="gpt-4.1-mini",
)
INBOUND = """From: dana@northwind.example
Subject: pricing for about 80 seats
We are comparing three vendors this quarter and need a number for 80 seats
before our board meeting on the 14th. What does that cost?
"""
async def main(cap: int = 8) -> None:
try:
result = await Runner.run(triage, INBOUND, max_turns=cap)
except MaxTurnsExceeded as exc:
print(f"STOPPED by cap: {exc}")
return
print("--- path ---")
for item in result.new_items:
print(f"{item.agent.name:>14} {item.type}")
print("\n--- reply ---")
print(result.final_output)
tokens_in = tokens_out = 0
for response in result.raw_responses:
usage = getattr(response, "usage", None)
if usage:
tokens_in += usage.input_tokens
tokens_out += usage.output_tokens
print(f"\nlast agent: {result.last_agent.name}")
print(
f"model calls: {len(result.raw_responses)} "
f"input tokens: {tokens_in} output tokens: {tokens_out}"
)
if __name__ == "__main__":
cap = int(sys.argv[1]) if len(sys.argv) > 1 else 8
asyncio.run(main(cap))
For step 7, the deterministic router that replaces the triage agent:
keywords = ("pricing", "seats", "plan", "renewal")
start = enrichment if any(k in INBOUND.lower() for k in keywords) else reply_drafter
result = await Runner.run(start, INBOUND, max_turns=cap)
Your reply wording will differ run to run, and the exact ordering of items shifts if the model requests both tools in a single turn. What should not differ: two handoff items in the default version, one in the step-7 version, and last agent: Reply Drafter in both.
The cap is the system, not a safety feature
max_turns=8 is the line that makes this thing shippable. The Agents SDK counts one turn per model call and raises MaxTurnsExceeded when the loop exceeds the limit, which means the ceiling is enforced in Python, outside the model, where no amount of clever prompting can talk its way past it. The SDK's own default is max_turns=10; 8 is deliberately tighter, because the happy path here is five model calls at most — triage, enrichment, one or two tool returns, drafter — and anything past that is a loop rather than work.
Unbounded loops are the most common way multi-agent systems fail in production, and the mechanism is boring. Agent A is unsure, so it hands to agent B. Agent B lacks the information it needs, so it hands back. Neither agent is broken. The graph is. You can reproduce it in this file by adding reply_drafter.handoffs = [enrichment] before main() and loosening the drafter's final instruction — then watch the model calls: count climb until the cap fires.
Three defences hold, in the order I add them. First, a directed graph with exactly one terminal agent: in the code above reply_drafter has no handoffs at all, so it cannot pass the ball back because it has nowhere to pass it. Second, a turn cap on every run — a parameter, not a prompt instruction, because the prompt is the thing that failed. Third, a per-agent tool allowlist: enrichment holds both tools and neither of the other two agents can call anything, so the agent that drafts customer-facing text has no write access by construction. A related SDK default is worth knowing while you debug tool loops: Agent.reset_tool_choice defaults to True, which clears a forced tool_choice="required" after the first tool call so the setting cannot spin forever. The layer after those three is guardrails — the SDK runs input guardrails on the first agent and output guardrails on the last, and a tripped one raises InputGuardrailTripwireTriggered or OutputGuardrailTripwireTriggered before any text is returned.
What I would not delegate at this stage: spend authority, irreversible writes, and anything that sends to a real human without a review gate. Not because agents are incapable — because the failure mode of a mis-specified graph is "does the wrong thing four times fast," and you want that to be cheap. I went through which GTM steps survive automation and which ones I hold back in agentic go-to-market, and the parts I would not automate.
The token ledger at the end of main() exists for the same reason. A multi-agent run costs more than a single call by a multiple you should measure on your own traffic rather than estimate. result.raw_responses carries one usage object per model call with input_tokens, output_tokens and total_tokens, and the SDK also aggregates the whole run in result.context_wrapper.usage, including a requests count. Metering before scaling is the standard discipline for anything billed per token, and the same argument applies to products, as set out in build a SaaS app with AI, starting with the token meter.
When it doesn't work
No API key set. The openai client raises on the first model call: openai.OpenAIError: The api_key client option must be set either by passing api_key to the client or by setting the OPENAI_API_KEY environment variable. A tracing warning about a missing key usually shows up first. Fix: export the key in the same shell you run python from, and re-check with the one-liner in step 2.
You hit the cap on a normal run. MaxTokensExceeded is not it — the exception is agents.exceptions.MaxTurnsExceeded, and the message names the limit you set. If 8 turns is not enough for three agents and two tools, the cap is not the problem. It is an agent that keeps re-calling a tool because its instruction does not say when to stop, or two agents that can each hand to the other. Print result.new_items types, find the repeating pair, remove one edge from the graph. Raising the number just buys the loop more runway.
asyncio.run() inside a notebook. RuntimeError: asyncio.run() cannot be called from a running event loop. Jupyter already owns the loop. Fix: call await main() directly in the cell, or use the synchronous wrapper Runner.run_sync(triage, INBOUND, max_turns=8) instead.
An unknown model name. openai.NotFoundError with a message that the model does not exist or you do not have access to it. Model identifiers change; check the model list on your own account rather than copying a string from an article, including this one.
What to change next
Swap the handoffs= chain for enrichment.as_tool(tool_name="enrich", tool_description="...") on a single orchestrator agent and run the same input. as_tool() hands back the sub-agent's final output as a string by default, and takes a custom_output_extractor= callable if you need to return something narrower than the whole last message. Then read both traces side by side. The handoff version shows control moving between agents with the conversation travelling with it; the tool version keeps one agent in charge and folds each sub-agent's answer back into one growing context. One of those two shapes fits your problem and the other will annoy you for a month, and twenty minutes of reading traces tells you which — a faster answer than any framework comparison, including the ones I have written about self-hosted agent builders.
Sources
- OpenAI Agents SDK — documentation — Agent/Runner primitives, installation, Python package name openai-agents
- OpenAI Agents SDK — Handoffs — handoffs= parameter, handoff_description, how a handoff is exposed to the model as a tool
- OpenAI Agents SDK — Running agents — max_turns, MaxTurnsExceeded, result.final_output, result.last_agent, result.new_items, raw_responses
- OpenAI Agents SDK — Tools — @function_tool schema generation from type hints and docstrings; agent.as_tool() for the supervisor pattern
- OpenAI Agents SDK — Tracing — tracing is on by default and traces are viewable in the OpenAI dashboard
- openai/openai-agents-python on GitHub — source, release history, supported Python versions
- OpenAI API keys — where to create the OPENAI_API_KEY used in step 2
- OpenAI API pricing — current per-token prices for the models named in the code
- Anthropic — How we built our multi-agent research system — reported that multi-agent systems used roughly 15x more tokens than chat interactions
- Cognition — Don't Build Multi-Agents — argument that split context across parallel agents produces incoherent results
- LangGraph — alternative framework for the same topology, graph-first instead of handoff-first