What an AI agent framework actually is
A library that owns the agent loop, the state between steps, and the stop condition — not a prompt wrapper, not a workflow engine, not MCP.
- AI agent framework
- An AI agent framework is a code library that owns an LLM's action loop: it runs the model, executes the tool calls it asks for, carries state between steps, and enforces when to stop.
The loop and the stop are the whole product. Prompt formatting, model switching and retries on 429 are commodity features that a dozen libraries do identically. What separates a framework from an HTTP client is that the model's output changes control flow at runtime, and something outside the model has to decide when that stops.
Where the term came from
The loop was described before the libraries existed. ReAct (Yao et al., arXiv, October 2022) formalised interleaving reasoning traces with tool calls, so a model could act, read an observation, and act again. AutoGPT, released March 2023, turned that into a program you could run on your laptop. It also made the signature failure famous: with no bound on iterations, the loop runs until the API key or the wallet gives out.
Nearly everything built since is a response to that. OpenAI shipped Swarm as an experimental multi-agent library in late 2024, archived it, and replaced it with the Agents SDK in March 2025. LangGraph models the agent as a graph with persisted state. CrewAI uses roles and crews. AutoGen uses conversations between agents. Pydantic AI uses types and validated outputs. Semantic Kernel uses plugins and planners.
Same loop. Different opinion about what shape the loop should take and where you are allowed to interrupt it.
Two different things get called an agent framework
The word covers two products that share almost no implementation.
| Framework-as-library | Framework-as-platform | |
|---|---|---|
| What you touch | A package: openai-agents, langgraph, crewai | A web UI with a canvas |
| Where the loop runs | Your process, your logs | The vendor's runtime |
| Version control | Pinned in a lockfile | Whatever the vendor deployed today |
| Who sees the raw tool calls | You | Whatever the dashboard shows |
| Named examples | LangGraph, CrewAI, AutoGen, Pydantic AI, Semantic Kernel, OpenAI Agents SDK | Hosted agent builders and no-code agent nodes |
Both are legitimate. They are not substitutes. If someone tells you they standardised on an agent framework, ask which of the two they mean, because the answer determines whether you can reproduce a run from six weeks ago. More on that trade-off in self-hosted AI agent builders.
What it is not
Not a prompt or model-wrapper library. A library that formats messages, normalises providers and retries on 429 is useful, and it is not a framework by this definition. It does not own the loop. You still write the while.
Not a workflow engine. Anthropic's distinction is the cleanest one in print: in a workflow, control flow is fixed by the author; in an agent, the model decides the path at runtime. Airflow, Temporal and a branching automation canvas are workflow engines. Many "agents" in production are workflows, and that is the better engineering choice, because a fixed path is testable and a model-chosen path is not.
Not MCP. The Model Context Protocol, announced November 2024, is a wire protocol for exposing tools and data to models. It standardises the plug. The framework is what calls it. Frameworks consume MCP servers; MCP does not run a loop, hold conversational state, or cap your spend.
Not an agent. The framework is the harness. The agent is a specific configuration of instructions, tools and limits running inside it.
A worked example, in about ten lines
from agents import Agent, Runner, function_tool
@function_tool
def lookup_invoice(account_id: str) -> dict:
return billing.get(account_id) # your code, not the model's
triage = Agent(
name="Billing triage",
instructions="Answer from the invoice only. If the invoice is missing, hand off.",
tools=[lookup_invoice],
)
result = await Runner.run(triage, "Why was account 8812 charged twice?", max_turns=6)
Four things the framework did that you would otherwise hand-write: it generated the JSON tool schema from the Python type hints, it ran the model, it executed lookup_invoice when the model asked for it and appended the result to state, and it raised MaxTurnsExceeded at turn seven instead of looping forever. LangGraph does the equivalent with a recursion limit that raises GraphRecursionError.
One thing it did not do: decide whether a double charge gets refunded.
Four tests for whether something qualifies
- Does it run the loop, or do you? If your code contains the iteration, it is a client library.
- Does state survive a step? The observation from turn three must be visible at turn four without you threading it manually. A framework that forgets is a prompt template.
- Can the model change the path at runtime? If every branch is drawn at author time, you have a workflow engine. Fine thing to own. Different thing.
- Is there a stop enforced outside the model? A turn ceiling like
max_turns=6, a recursion limit, a budget, a wall clock. Unbounded iteration is the most commonly documented production failure for agents, and no amount of instruction text fixes it, because the instruction text is evaluated by the thing that is looping.
Three of four is not a pass. Test four is the one that costs real money when it fails.
The part the framework will never give you
Every framework ships a cap and a default. The default is the vendor's guess, not your risk tolerance. The cap that matters is the one tied to a dollar figure and an owner, and that lives in your code and your contract, not in the library. I would not hand an agent unsupervised spend authority on the strength of a max_turns argument.
Choosing between LangGraph, CrewAI and the Agents SDK is a half-day decision, and all three will run the loop correctly. Deciding which actions require a human signature is the harder one, and it does not get easier with a better framework — see three hard caps and a three-agent system that stops when it should.
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al.) — October 2022 paper that formalised the reason-act-observe loop every agent framework implements
- AutoGPT repository — March 2023 project that shipped the loop as a runnable program and popularised the unbounded-loop failure
- Anthropic — Building effective agents — Workflow vs agent distinction: control flow fixed at author time vs decided by the model at runtime
- Introducing the Model Context Protocol — MCP announced November 2024 as a protocol for tool exposure, not a framework
- Model Context Protocol docs — Spec and server/client roles that frameworks consume
- OpenAI Agents SDK documentation — Runner, function_tool, handoffs, max_turns and MaxTurnsExceeded; tool schemas generated from type hints
- OpenAI Swarm repository (archived) — Experimental multi-agent library, since replaced by the Agents SDK
- LangGraph repository — Graph-shaped framework with persisted state and a recursion limit that raises GraphRecursionError
- CrewAI repository — Role-and-crew shaped framework
- Microsoft AutoGen repository — Conversation-shaped multi-agent framework
- Pydantic AI documentation — Type-validated agent framework with structured outputs
- Semantic Kernel repository — Microsoft's plugin/planner-based framework