Gil Allouche
← All writing
ReferenceAutonomous agents

Self-hosted AI agents: what to run yourself, what to rent

A working definition, the four layers you can host, the VRAM arithmetic, the license traps, and the layer I would not self-host first.

October 10, 2026·Gil Allouche·10 min read
A reference explainer. Every factual claim links to its source; where this is an opinion from operating experience, it says so.

The weights stopped being the hard part

OpenAI shipped gpt-oss-120b and gpt-oss-20b under Apache 2.0, with the 20b model sized to run in 16 GB of memory and the 120b on a single 80 GB GPU (OpenAI). Pull it with Ollama, which is MIT and serves an HTTP API on port 11434 (repo). That part is a Saturday.

A self-hosted AI agent is one whose loop — planning, tool calls, memory, retries, and the caps that stop it — runs on infrastructure you control. Whether you also run the model weights is a separate decision, and usually the harder one.

Most people asking mean both at once. Two projects, two bills, two failure modes. Separate them before you buy a GPU.

Four layers, and the one I would rent first

An agent is not one thing you host. It is four, and you can draw the line anywhere.

Which layers of a self-hosted agent stack to host and which to rentSelf-hostRentAgent loop + capsTool calls + credsMemory + stateModel inferenceMove left later

The loop, the tools, and the state store are where your actual product lives. Those belong to you on day one. Inference is a commodity, and swapping it is a config change if you built against an OpenAI-compatible endpoint — which vLLM serves natively (repo), and which llama.cpp's server speaks too (repo).

Host the loop, rent the tokens, keep the inference layer swappable. Move inference in-house when a specific constraint forces it: data that cannot leave your network, a per-token bill that exceeds a reserved GPU at your volume, or latency on a small model you call thousands of times an hour. Nothing else counts as a reason.

The stack people actually run

LayerCommon self-hosted choiceLicense to read firstDominant failure
InferencevLLM (Apache 2.0), Ollama (MIT), llama.cpp (MIT)Permissive — the weights are the trap, not the serverIdle GPU hours; OOM under concurrency
OrchestrationLangGraph (library is MIT), CrewAI, AutoGenLibrary vs platform split — LangChain's platform tiers are paidUnbounded loops
Workflow surfacen8n, DifySustainable Use License; Apache 2.0 with added conditionsLicense surprise when you resell
ToolsMCP servers, Playwright in DockerPer-server, varies wildlyCredential blast radius
StatePostgres + langgraph-checkpoint-postgresPostgreSQL LicenseRuns lost on restart

n8n is source-available under its Sustainable Use License, not OSI open source: internal self-hosting is fine, and hosting it as a service that competes with n8n is not (docs). Dify is Apache 2.0 with additional conditions covering multi-tenant service use and branding (LICENSE). Neither is a problem for an internal agent. Both are a problem if your plan is to white-label the thing and sell it, and that discovery is free on day one and expensive at the term sheet.

The weights are the other trap. "Open" in a launch headline often means a custom community license with acceptable-use terms and a gated download, which is how Llama 3.3 70B Instruct is distributed (model card). gpt-oss is genuinely Apache 2.0 (OpenAI). Different legal objects. Read the file, not the blog post.

VRAM arithmetic, before you pick a box

Model weights at 4-bit quantization need roughly half a byte per parameter. A 70B model is therefore about 35 GB of weights — arithmetic, not a benchmark — and that is before the KV cache, which grows with context length and with the number of concurrent requests you serve (vLLM docs).

This is the number people get wrong. They size the GPU for the weights, test with one request, then watch it fall over when four agents run in parallel with 40k-token contexts. Agents are a concurrency workload with long contexts. That is the worst case for KV cache pressure.

An AWS g5.xlarge carries a single NVIDIA A10G with 24 GB of GPU memory (instance specs). That holds a quantized 20B-class model with room for a modest cache. It does not hold a 70B at useful context with concurrency. gpt-oss-120b is specified for a single 80 GB GPU (OpenAI) — a different class of machine, and a different conversation with finance.

The serving engine matters as much as the card. vLLM is the choice once you have concurrent agents and a datacenter GPU: continuous batching, PagedAttention, OpenAI-compatible server (repo). Ollama gets you a working endpoint on 11434 in minutes and is the right tool for development and single-user agents (repo) — putting it behind a production multi-agent loop is the common mistake, because the concurrency you are about to throw at it is not what it is for. llama.cpp and GGUF quantization are for the boxes with no datacenter GPU at all (repo).

If you are still deciding what the orchestration layer above this even does, what an AI agent framework actually is covers the division of labour.

What breaks, and where the fix has to live

The documented production failures of self-hosted agents are not exotic. The same four, every time.

Unbounded loops. The agent retries, re-plans, re-reads the same page, and burns tokens until something external stops it. The fix is a hard step cap and a wall-clock timeout enforced outside the model — in the orchestrator, in code, as a counter the model cannot see or argue with. A prompt that says "do not loop" is not a cap. The three-cap pattern is spelled out in one agent, three hard caps, and the stopping logic in build a three-agent system that stops when it should.

Prompt injection through tool output. Anything your agent reads — a web page, a ticket, an email, a repo README — is untrusted input that can carry instructions. OWASP ranks prompt injection as LLM01 in its Top 10 for LLM Applications (OWASP). Self-hosting does not reduce this risk by one percent. It changes who is on call for it.

Excessive agency. OWASP's LLM06 (same list) is the agent having more permission than the task requires: a read-only research agent holding a write token, a support agent with refund authority it was never meant to use. Self-hosted stacks make this worse by default, because environment variables are easy and scoped credentials are work.

Lost state. An agent that keeps its plan in process memory loses the run when the container restarts, and a mid-flight run that restarts from the top will repeat side effects it already committed. Checkpoint to Postgres and make tool calls idempotent (langgraph-checkpoint-postgres).

The credential, not the GPU, is the thing that will hurt you. A self-hosted agent runs with long-lived environment-variable secrets in a container that also executes model-chosen tool calls on untrusted text — which OWASP catalogues as prompt injection feeding excessive agency. Scope every token to one resource, set an egress allowlist, and give the agent nothing with spend or delete authority.

On the MCP side: the 2025-06-18 revision of the Model Context Protocol specification sets out the transport and authorization requirements for servers, and remote HTTP servers are expected to authorize requests rather than trust the caller (spec). Read that section before you decide the VPC boundary is your authorization layer. It is not.

The bill has a different shape, not just a different size

Rented inference bills per token. Self-hosted inference bills per hour, awake or idle. That structural difference decides the answer more often than any price comparison does.

An agent with bursty traffic — 200 runs between 9am and 11am, nothing overnight — pays a GPU to sit idle for 22 hours a day. Steady high volume on a small model is the opposite case, and the only one where moving inference in-house pays for itself.

So work out your token volume first. A support-triage agent: 12 model turns per run, roughly 8,000 input tokens and 600 output tokens per turn, 3,000 runs a month.

What would this cost you?
40
15,000
5,000
per run
$2.70
full pass
$13,500

Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.

Input tokens dominate, because every turn re-sends the growing transcript. Turn count is the most expensive variable you control, which is why the step cap is a cost control before it is a safety control. And 3,000 runs a month sounds small — a hundred a day — right up until 12 turns makes it 36,000 model calls.

Then price the self-hosted side honestly: GPU hours for the full month, not the hours you expect traffic, plus storage for weights, plus the engineer-days to keep vLLM and a driver stack current. If the per-token number is smaller than that, rent.

What I would not do

I would not self-host inference as the first move. It is the layer with the most operational cost and the least product differentiation, and it is the one people reach for first because it feels like the real work. Hosting a GPU does not make your agent better at its job. It makes you responsible for CUDA versions.

I would not use a visual workflow builder as the control plane for anything with spend or write authority. An n8n or Dify canvas is good for wiring and good for demos. The caps, the idempotency, and the audit trail want to be code you can test, and a canvas hides exactly the control flow you most need to read at 3am.

I would not treat self-hosting as a privacy answer while still calling a hosted model API. A stack where the orchestrator sits in your VPC and every prompt goes out to a third-party endpoint has the operational burden of self-hosting and the data posture of SaaS — the worst trade available. Pick one deliberately. If the requirement is genuinely that data cannot leave, inference comes in-house too, and the GPU bill is the price of that requirement rather than an optimisation.

And I would not give an autonomous agent spend authority yet, self-hosted or not. That is a judgement, not a benchmark. The gap between "the agent chose a reasonable action" and "the agent chose a reasonable action on every input an adversary can construct" is still wide enough that I want a human on the payment. Starting a business with AI agents, and what stays human is where I draw that line in more detail.

The version you pin is the version you own

This gets decided by accident. When you self-host weights, nobody deprecates your model — and nobody improves it either. A hosted endpoint silently gets better and occasionally silently changes behaviour; a local checkpoint does exactly what it did in March, forever, until you do the upgrade work yourself.

Determinism is worth real money here. An evaluation suite that passed on a pinned local model keeps passing, and a prompt you tuned against specific tokenizer behaviour stays tuned. Agents are brittle in proportion to how much their prompts encode the quirks of one model, so pinning removes a whole class of Monday-morning regressions.

The cost arrives later, as drift. Tool-calling reliability and long-context behaviour are where open-weight models have been moving fastest, and those are precisely the capabilities an agent loop depends on. A stack frozen two releases back is paying GPU hours to make worse tool calls than a current hosted endpoint would make for the token price.

So budget the upgrade. Keep an evaluation set of 50–100 real runs with known-good outcomes, re-run it against each candidate model, and treat a model swap as a release with a rollback, not a config tweak. If you have not built that evaluation set, you have not finished self-hosting. You have moved the inference and kept the uncertainty.

Sources

  1. Introducing gpt-oss (OpenAI) — Apache 2.0 open-weight models; gpt-oss-20b sized for 16 GB, gpt-oss-120b for a single 80 GB GPU
  2. gpt-oss-20b model card — Weights, license and memory footprint of the smaller open-weight model
  3. Ollama repository — MIT license, local HTTP API, default port 11434
  4. vLLM repository — Apache 2.0 inference server with an OpenAI-compatible API
  5. llama.cpp repository — MIT license, GGUF quantized inference on CPU and consumer GPUs
  6. LangGraph repository — MIT-licensed orchestration library, separate from the commercial platform
  7. LangChain pricing — The managed and self-hosted platform tiers around the open library are paid
  8. langgraph-checkpoint-postgres — Durable agent state in Postgres rather than process memory
  9. Dify license — Apache 2.0 with additional conditions on multi-tenant service and branding
  10. Llama 3.3 70B Instruct model card — Gated weights under a custom community license, not MIT or Apache
  11. Model Context Protocol specification, revision 2025-06-18 — Transport and authorization requirements for MCP servers
  12. OWASP Top 10 for LLM Applications — Prompt injection and excessive agency as documented top risks
  13. Amazon EC2 G5 instances — g5.xlarge carries one NVIDIA A10G with 24 GB of GPU memory
  14. Playwright Docker image — Official container for running browser automation tools in isolation
  15. vLLM documentation — PagedAttention and KV-cache behaviour under concurrency

Related