Gil Allouche
← All writing
ReferenceAutonomous agents

Self-hosted AI agent builders: what you actually control

A reference on self-hosted agent orchestration — the licenses, the four layers you can host independently, and the one layer I would not self-host first.

September 28, 2026·Gil Allouche·13 min read
A reference explainer. Every factual claim links to its source; where this is an opinion from operating experience, it says so.

The answer, and the four layers people bundle together

A self-hosted AI agent builder is an orchestration control plane you run on your own infrastructure — n8n, Dify, Flowise, Langflow, or a framework like LangGraph or CrewAI sitting in your own repo. Self-host the orchestrator first. Not the model.

n8n ships under the Sustainable Use License: source-available, not OSI open source. You can run it internally for free, forever, commercial use included. You cannot resell it as a hosted service without a separate agreement. Dify does something similar — Apache 2.0 plus additional conditions covering multi-tenant service use and frontend branding.

"Self-hosted" is four separate decisions that most comparison posts mash into one:

  1. The builder surface. The canvas, the prompt editor, the credential vault.
  2. The orchestrator and its state. What picks the next step, and where the run history lives — Postgres, in practice.
  3. The tool layer. HTTP calls, database queries, MCP servers — the Model Context Protocol Anthropic published in November 2024 and since adopted by OpenAI's Agents SDK — anything with side effects.
  4. Inference. The model weights and the GPUs under them.

You can host any subset. Most teams that say "we need self-hosted agents" host layers 1 through 3 and call a model API for layer 4. That is the right default. It is also the exact configuration that does not satisfy a contract saying customer data may not leave your environment.

The four layers of a self-hosted agent stack and the egress boundaryYour infrastructure1. Builder canvas2. Orchestrator + Postgres3. Tools / MCP servers4. vLLM on GPUs (optional)Model APIegressprompts, tooloutput, PII

The licenses are the first filter

Before you compare node counts, read the LICENSE file. Two products that both say "self-hosted, open source" on the homepage can put you in completely different legal positions, and the difference bites the moment you want to embed one in something you sell.

ProjectLicense as published in the repoWhat that means for you
n8nSustainable Use License (fair-code)Internal and commercial self-hosting fine; hosting it as a service for others is not
DifyApache 2.0 + additional conditionsSelf-host freely; multi-tenant service use and removing frontend branding are restricted
LangflowMITNo restrictions worth planning around
FlowiseApache 2.0Permissive; check the current file before you embed it
CrewAIMITLibrary in your repo, no control plane to license
LangGraph (library)MITFree; self-hosted LangGraph Platform deployment is an Enterprise-plan product
TemporalMITDurable execution engine, fully self-hostable
Open WebUIRevised BSD-3-Clause with a branding-retention conditionFine internally; read it before white-labelling

That LangGraph row is the one that catches teams. The graph library is MIT and runs anywhere Python runs. The managed control plane around it — the persistence layer, the task queue, the deployment UI — is a separate commercial product, and self-hosted deployment of it sits on the Enterprise plan. So if your plan was "self-host LangGraph Platform to avoid the vendor," you are buying a vendor relationship anyway. You are just also running the servers.

Licenses drift, and the record is public: Elastic moved Elasticsearch to SSPL in 2021, HashiCorp put Terraform under BSL 1.1 in August 2023, Redis moved to RSALv2/SSPL in March 2024 and then to AGPLv3 with Redis 8 in May 2025, and n8n itself came to the Sustainable Use License from Apache 2.0 with the Commons Clause. Source-available relicensing has hit enough infrastructure projects that the current LICENSE file is best treated as a snapshot of vendor intent rather than a promise. The hedge is boring: keep your agent's actual logic — prompts, tool definitions, control flow — in your own repository in your own language, so swapping the surface around it is a week of work rather than a rewrite.

Where the drag-and-drop canvas stops being the right container

A visual builder is excellent for the first two weeks and structurally wrong as the permanent home for an agent that runs unattended.

n8n stores workflows as JSON. Langflow exports flows as JSON. Dify exports apps as a YAML DSL. All three are technically diffable, and none of them produce a diff a human can review. n8n writes a position: [x, y] pair and a generated ID onto every node; Langflow's export carries the same canvas coordinates alongside each component's template. Change one prompt in a fifteen-node graph and your pull request is a few hundred lines of reordered position coordinates and regenerated node IDs. Now answer the question that actually matters after an incident: which change made the agent start doing that?

The second problem is testing. An agent that decides its own next step — the definition I use — has a branching factor, and you cannot assert on behaviour you cannot invoke from a test runner. Three candidate tools at each of five decision points is already 3⁵ = 243 distinct paths through one run. Canvas tools give you a "run once" button and an execution log. That is a debugger, not a test suite.

So: canvas for the plumbing, code for the decisions. Triggers, webhooks, retries, credential storage, the 400-odd integrations you do not want to write — n8n's queue mode with Postgres and Redis handles that well, and the native AI Agent and LLM chain nodes get a prototype running in an afternoon. The loop that chooses tools, the step budget, the stopping condition, the escalation path all live somewhere else: in a repo, behind tests, with a version number.

Self-host orchestrator, rent the model

Inference is the layer I would self-host last

Do not buy or reserve GPUs before you know the token volume. Self-hosted inference is not the problem — vLLM is genuinely good, Apache 2.0, came out of UC Berkeley's Sky Computing Lab, and its continuous batching and PagedAttention KV cache management are why it beats a naive server by a wide margin: the SOSP 2023 PagedAttention paper reports up to 24× the throughput of a HuggingFace Transformers server and holds KV cache waste under 4%, against the 60–80% waste typical of contiguous pre-allocation. The problem is that this is a capacity bet, and on day one you have no data.

The shape of the mistake is predictable. You size a cluster for peak, the agent turns out to fire 200 times a day instead of 20,000, and you have converted a small variable cost into a fixed one that is not small. The fixed side is checkable before you commit: rented 80 GB H100 capacity is commonly quoted in the $2–3/hour range, so one card left running is roughly $1,500–2,200 a month before storage and egress, and the two-card node an FP16 70B needs is double that. Rented hourly GPUs soften this. They do not remove it, because the moment you stop the instance to save money, your agent's tool latency goes from 800 ms to a cold start — pulling ~140 GB of weights off local NVMe and into HBM is minutes, not seconds.

Ollama is MIT-licensed and installs in one command, which makes it right for local development on a single node and wrong for concurrent production traffic. vLLM is the production answer. The ordering:

  1. Ship the agent against a hosted model API. Note the exact model string you used, because behaviour changes between versions.
  2. Put a self-hosted litellm proxy in front of it from day one — MIT, OpenAI-compatible across 100+ providers, one endpoint, per-key spend tracking and budgets. This is the swap point later.
  3. Run for a few weeks. Record tokens in, tokens out, and calls per run. Instrument it like a product, because this is the meter that makes the next decision for you.
  4. Compare the measured monthly token bill against a rented GPU of the size your model actually needs — the ~$1,800-a-month single card, or twice that for an unquantised 70B. If the API bill is not comfortably larger, stop here.
  5. Only then stand up vllm serve behind the same proxy and shift traffic by key. Your agent code does not change.

Step 2 is the load-bearing one. A gateway you control turns "self-host the model" from an architecture decision into a config change, which is the only state in which anyone makes it calmly.

The failure modes that arrive with the server

Managed platforms hide a set of problems. Self-hosting hands them to you, and every one of these is documented well enough that you should not have to discover it in production.

Unbounded loops. The most common production failure in agent systems is an agent that keeps calling tools without converging. The fix is a step cap enforced outside the model, in the runtime, not requested in the prompt. LangGraph does this by default: a recursion_limit of 25 super-steps, after which the graph raises GraphRecursionError rather than continuing. CrewAI exposes the equivalent knobs explicitly — max_iter per agent, max_rpm on the crew. If your builder has no equivalent, you are relying on a language model to decide when it has done enough. That is the same as having no cap.

Secrets in the export. Canvas tools separate credentials from workflow JSON — n8n encrypts its credential store with the value in N8N_ENCRYPTION_KEY — and then someone pastes an API key into an HTTP node header as plain text. It is now in your Git history. Credentials belong in the platform's vault or an external secret manager, and the workflow export belongs in review.

Webhooks facing the internet. A self-hosted builder with a public trigger URL is an unauthenticated RPC endpoint into your tool layer. n8n's Webhook node offers basic-auth and header-auth options, and its docs push you to a reverse proxy with TLS in front of both; take the push.

Single-process scaling. n8n on SQLite with one process is a prototype. Production is EXECUTIONS_MODE=queue with Postgres, Redis and separate worker processes each carrying their own --concurrency setting, and moving to that later is a migration, not a flag.

No replay. Almost no self-hosted builder ships trace-level replay or evaluation out of the box. If you cannot re-run yesterday's failing execution against today's prompt, you are debugging by anecdote. Temporal's durable execution model is the heavyweight fix — MIT, self-hostable, a descendant of Uber's Cadence, and it replays a workflow's recorded event history to rebuild state after a crash, which makes a long-running agent resumable instead of restarted.

A step cap in the system prompt is not a step cap. Enforce the budget in the runtime that calls the model, where the model cannot talk its way past it.

If "the data cannot leave" is your reason, check the egress

This is where I see the most expensive confusion. A team self-hosts n8n specifically because of a data residency requirement, then wires an OpenAI node into it. The orchestrator is in the VPC. The customer records are in the request body going out over TLS to someone else's inference cluster.

That is defensible on paper. OpenAI's documented position is that API inputs and outputs are not used to train models by default, with retention of up to 30 days for abuse monitoring and zero-data-retention available to eligible customers. The in-cloud managed services carry their own version of the same clause: Azure OpenAI also logs prompts and completions for up to 30 days for abuse monitoring unless a customer is approved for the modified abuse-monitoring exemption. But "acceptable under a documented policy" and "the data never leaves our environment" are different claims, and only one of them survives a security questionnaire. Decide which one you are making before you design the stack around it.

ConfigurationWhat leavesWhat you can actually claim
Self-hosted orchestrator, public model APIPrompts, tool output, retrieved recordsCovered by the provider's terms, not by your network boundary
Self-hosted orchestrator, in-cloud managed inference (Bedrock, Vertex, Azure OpenAI, your own region and account)Nothing outside your cloud tenancy and regionSatisfies a residency requirement, at a fraction of the work of GPUs
Fully self-hosted weights on your own vLLMNothingNothing leaves — and you own model upgrades, capacity and the GPU bill

The middle row is the one teams skip, and it is the right answer more often than either of its neighbours. It clears the compliance constraint without making you an inference operator.

What I would run

Concretely, for a team that already knows the basics and wants an agent in production this month:

Orchestration in a repo, not a canvas — LangGraph if you want an opinionated graph, CrewAI if the work decomposes cleanly into roles, plain code if neither fits. State in Postgres, through LangGraph's PostgresSaver checkpointer or your own tables. Tools behind MCP servers, because the protocol lets you run the tool layer as its own deployable with its own auth instead of as nodes glued into the workflow. A self-hosted LiteLLM proxy as the only thing that talks to a model, so the provider is a config value. A hard step budget in the 10–20 tool-call range, enforced in the runtime. Keep n8n or Dify around for triggers, connectors and the human-facing pieces; rewriting 400 integrations is not a good use of anyone's month. Temporal the moment a single agent run exceeds a few minutes.

What I would not use: a drag-and-drop canvas as the source of truth for an autonomous agent. The tools are not weak — n8n's AI nodes are good and Dify's app model is well thought out. But an agent that runs unattended needs a diff, a test and a version, and a JSON graph gives you none of the three in a form a human can review at 2am.

And I would not self-host inference in the first month. AI doing the work by default is a claim about who performs the task, not about where the weights sit. Running your own GPUs is a cost optimisation with a compliance side effect, and optimisations you make before you have the meter reading are guesses with a monthly invoice attached.

Self-hosting does not reduce the number of things that can break. It moves them onto your pager, in exchange for a boundary you can point at in a contract.

Sources

  1. n8n repository — Primary source for n8n's license file and codebase
  2. n8n queue mode docs — Scaling n8n past a single process requires Postgres plus Redis and separate worker processes
  3. n8n advanced AI docs — n8n's native AI Agent and LLM chain nodes
  4. Dify LICENSE — Dify is Apache 2.0 with additional conditions covering multi-tenant service use and frontend branding
  5. Dify documentation — Dify apps export to a YAML DSL file
  6. Langflow repository — Langflow's MIT license and flow-as-JSON export
  7. Flowise repository — Flowise license and self-hosting instructions
  8. CrewAI repository — CrewAI is MIT licensed
  9. LangGraph repository — The LangGraph library itself is MIT licensed
  10. LangChain pricing — Self-hosted LangGraph Platform deployment sits on the Enterprise plan, unlike the MIT library
  11. LangGraph recursion limit error — LangGraph enforces a default recursion_limit of 25 super-steps outside the model
  12. Temporal repository — Temporal server is MIT licensed and self-hostable for durable execution
  13. Open WebUI repository — Open WebUI's license is a revised BSD-3-Clause with a branding-retention condition
  14. vLLM documentation — vLLM's Apache 2.0 self-hosted inference server, continuous batching and PagedAttention KV cache management
  15. Ollama repository — MIT-licensed local model runner used for single-node self-hosted inference
  16. Llama 3.3 70B Instruct model card — 70 billion parameters and a 128k context window — the basis for the VRAM arithmetic
  17. NVIDIA H100 product page — 80 GB HBM per H100 SXM card
  18. OpenAI API data controls — API inputs and outputs are not used for training by default; retention window and zero-data-retention options
  19. OpenAI enterprise privacy — Retention and ZDR commitments for API traffic
  20. Model Context Protocol — Open protocol for exposing tools to agents; lets the tool layer be hosted separately from the orchestrator
  21. LiteLLM repository — MIT-licensed self-hosted proxy that gives one OpenAI-compatible endpoint across providers, with per-key spend tracking

Related