Self-hosted AI agent builders: what you actually control
A reference on self-hosted agent orchestration — the licenses, the four layers you can host independently, and the one layer I would not self-host first.
The answer, and the four layers people bundle together
A self-hosted AI agent builder is an orchestration control plane you run on your own infrastructure — n8n, Dify, Flowise, Langflow, or a framework like LangGraph or CrewAI sitting in your own repo. Self-host the orchestrator first. Not the model.
n8n ships under the Sustainable Use License: source-available, not OSI open source. You can run it internally for free, forever, commercial use included. You cannot resell it as a hosted service without a separate agreement. Dify does something similar — Apache 2.0 plus additional conditions covering multi-tenant service use and frontend branding.
"Self-hosted" is four separate decisions that most comparison posts mash into one:
- The builder surface. The canvas, the prompt editor, the credential vault.
- The orchestrator and its state. What picks the next step, and where the run history lives — Postgres, in practice.
- The tool layer. HTTP calls, database queries, MCP servers — the Model Context Protocol Anthropic published in November 2024 and since adopted by OpenAI's Agents SDK — anything with side effects.
- Inference. The model weights and the GPUs under them.
You can host any subset. Most teams that say "we need self-hosted agents" host layers 1 through 3 and call a model API for layer 4. That is the right default. It is also the exact configuration that does not satisfy a contract saying customer data may not leave your environment.
The licenses are the first filter
Before you compare node counts, read the LICENSE file. Two products that both say "self-hosted, open source" on the homepage can put you in completely different legal positions, and the difference bites the moment you want to embed one in something you sell.
| Project | License as published in the repo | What that means for you |
|---|---|---|
| n8n | Sustainable Use License (fair-code) | Internal and commercial self-hosting fine; hosting it as a service for others is not |
| Dify | Apache 2.0 + additional conditions | Self-host freely; multi-tenant service use and removing frontend branding are restricted |
| Langflow | MIT | No restrictions worth planning around |
| Flowise | Apache 2.0 | Permissive; check the current file before you embed it |
| CrewAI | MIT | Library in your repo, no control plane to license |
| LangGraph (library) | MIT | Free; self-hosted LangGraph Platform deployment is an Enterprise-plan product |
| Temporal | MIT | Durable execution engine, fully self-hostable |
| Open WebUI | Revised BSD-3-Clause with a branding-retention condition | Fine internally; read it before white-labelling |
That LangGraph row is the one that catches teams. The graph library is MIT and runs anywhere Python runs. The managed control plane around it — the persistence layer, the task queue, the deployment UI — is a separate commercial product, and self-hosted deployment of it sits on the Enterprise plan. So if your plan was "self-host LangGraph Platform to avoid the vendor," you are buying a vendor relationship anyway. You are just also running the servers.
Licenses drift, and the record is public: Elastic moved Elasticsearch to SSPL in 2021, HashiCorp put Terraform under BSL 1.1 in August 2023, Redis moved to RSALv2/SSPL in March 2024 and then to AGPLv3 with Redis 8 in May 2025, and n8n itself came to the Sustainable Use License from Apache 2.0 with the Commons Clause. Source-available relicensing has hit enough infrastructure projects that the current LICENSE file is best treated as a snapshot of vendor intent rather than a promise. The hedge is boring: keep your agent's actual logic — prompts, tool definitions, control flow — in your own repository in your own language, so swapping the surface around it is a week of work rather than a rewrite.
Where the drag-and-drop canvas stops being the right container
A visual builder is excellent for the first two weeks and structurally wrong as the permanent home for an agent that runs unattended.
n8n stores workflows as JSON. Langflow exports flows as JSON. Dify exports apps as a YAML DSL. All three are technically diffable, and none of them produce a diff a human can review. n8n writes a position: [x, y] pair and a generated ID onto every node; Langflow's export carries the same canvas coordinates alongside each component's template. Change one prompt in a fifteen-node graph and your pull request is a few hundred lines of reordered position coordinates and regenerated node IDs. Now answer the question that actually matters after an incident: which change made the agent start doing that?
The second problem is testing. An agent that decides its own next step — the definition I use — has a branching factor, and you cannot assert on behaviour you cannot invoke from a test runner. Three candidate tools at each of five decision points is already 3⁵ = 243 distinct paths through one run. Canvas tools give you a "run once" button and an execution log. That is a debugger, not a test suite.
So: canvas for the plumbing, code for the decisions. Triggers, webhooks, retries, credential storage, the 400-odd integrations you do not want to write — n8n's queue mode with Postgres and Redis handles that well, and the native AI Agent and LLM chain nodes get a prototype running in an afternoon. The loop that chooses tools, the step budget, the stopping condition, the escalation path all live somewhere else: in a repo, behind tests, with a version number.
Inference is the layer I would self-host last
Do not buy or reserve GPUs before you know the token volume. Self-hosted inference is not the problem — vLLM is genuinely good, Apache 2.0, came out of UC Berkeley's Sky Computing Lab, and its continuous batching and PagedAttention KV cache management are why it beats a naive server by a wide margin: the SOSP 2023 PagedAttention paper reports up to 24× the throughput of a HuggingFace Transformers server and holds KV cache waste under 4%, against the 60–80% waste typical of contiguous pre-allocation. The problem is that this is a capacity bet, and on day one you have no data.
The shape of the mistake is predictable. You size a cluster for peak, the agent turns out to fire 200 times a day instead of 20,000, and you have converted a small variable cost into a fixed one that is not small. The fixed side is checkable before you commit: rented 80 GB H100 capacity is commonly quoted in the $2–3/hour range, so one card left running is roughly $1,500–2,200 a month before storage and egress, and the two-card node an FP16 70B needs is double that. Rented hourly GPUs soften this. They do not remove it, because the moment you stop the instance to save money, your agent's tool latency goes from 800 ms to a cold start — pulling ~140 GB of weights off local NVMe and into HBM is minutes, not seconds.
Ollama is MIT-licensed and installs in one command, which makes it right for local development on a single node and wrong for concurrent production traffic. vLLM is the production answer. The ordering:
- Ship the agent against a hosted model API. Note the exact model string you used, because behaviour changes between versions.
- Put a self-hosted
litellmproxy in front of it from day one — MIT, OpenAI-compatible across 100+ providers, one endpoint, per-key spend tracking and budgets. This is the swap point later. - Run for a few weeks. Record tokens in, tokens out, and calls per run. Instrument it like a product, because this is the meter that makes the next decision for you.
- Compare the measured monthly token bill against a rented GPU of the size your model actually needs — the ~$1,800-a-month single card, or twice that for an unquantised 70B. If the API bill is not comfortably larger, stop here.
- Only then stand up
vllm servebehind the same proxy and shift traffic by key. Your agent code does not change.
Step 2 is the load-bearing one. A gateway you control turns "self-host the model" from an architecture decision into a config change, which is the only state in which anyone makes it calmly.
The failure modes that arrive with the server
Managed platforms hide a set of problems. Self-hosting hands them to you, and every one of these is documented well enough that you should not have to discover it in production.
Unbounded loops. The most common production failure in agent systems is an agent that keeps calling tools without converging. The fix is a step cap enforced outside the model, in the runtime, not requested in the prompt. LangGraph does this by default: a recursion_limit of 25 super-steps, after which the graph raises GraphRecursionError rather than continuing. CrewAI exposes the equivalent knobs explicitly — max_iter per agent, max_rpm on the crew. If your builder has no equivalent, you are relying on a language model to decide when it has done enough. That is the same as having no cap.
Secrets in the export. Canvas tools separate credentials from workflow JSON — n8n encrypts its credential store with the value in N8N_ENCRYPTION_KEY — and then someone pastes an API key into an HTTP node header as plain text. It is now in your Git history. Credentials belong in the platform's vault or an external secret manager, and the workflow export belongs in review.
Webhooks facing the internet. A self-hosted builder with a public trigger URL is an unauthenticated RPC endpoint into your tool layer. n8n's Webhook node offers basic-auth and header-auth options, and its docs push you to a reverse proxy with TLS in front of both; take the push.
Single-process scaling. n8n on SQLite with one process is a prototype. Production is EXECUTIONS_MODE=queue with Postgres, Redis and separate worker processes each carrying their own --concurrency setting, and moving to that later is a migration, not a flag.
No replay. Almost no self-hosted builder ships trace-level replay or evaluation out of the box. If you cannot re-run yesterday's failing execution against today's prompt, you are debugging by anecdote. Temporal's durable execution model is the heavyweight fix — MIT, self-hostable, a descendant of Uber's Cadence, and it replays a workflow's recorded event history to rebuild state after a crash, which makes a long-running agent resumable instead of restarted.
If "the data cannot leave" is your reason, check the egress
This is where I see the most expensive confusion. A team self-hosts n8n specifically because of a data residency requirement, then wires an OpenAI node into it. The orchestrator is in the VPC. The customer records are in the request body going out over TLS to someone else's inference cluster.
That is defensible on paper. OpenAI's documented position is that API inputs and outputs are not used to train models by default, with retention of up to 30 days for abuse monitoring and zero-data-retention available to eligible customers. The in-cloud managed services carry their own version of the same clause: Azure OpenAI also logs prompts and completions for up to 30 days for abuse monitoring unless a customer is approved for the modified abuse-monitoring exemption. But "acceptable under a documented policy" and "the data never leaves our environment" are different claims, and only one of them survives a security questionnaire. Decide which one you are making before you design the stack around it.
| Configuration | What leaves | What you can actually claim |
|---|---|---|
| Self-hosted orchestrator, public model API | Prompts, tool output, retrieved records | Covered by the provider's terms, not by your network boundary |
| Self-hosted orchestrator, in-cloud managed inference (Bedrock, Vertex, Azure OpenAI, your own region and account) | Nothing outside your cloud tenancy and region | Satisfies a residency requirement, at a fraction of the work of GPUs |
| Fully self-hosted weights on your own vLLM | Nothing | Nothing leaves — and you own model upgrades, capacity and the GPU bill |
The middle row is the one teams skip, and it is the right answer more often than either of its neighbours. It clears the compliance constraint without making you an inference operator.
What I would run
Concretely, for a team that already knows the basics and wants an agent in production this month:
Orchestration in a repo, not a canvas — LangGraph if you want an opinionated graph, CrewAI if the work decomposes cleanly into roles, plain code if neither fits. State in Postgres, through LangGraph's PostgresSaver checkpointer or your own tables. Tools behind MCP servers, because the protocol lets you run the tool layer as its own deployable with its own auth instead of as nodes glued into the workflow. A self-hosted LiteLLM proxy as the only thing that talks to a model, so the provider is a config value. A hard step budget in the 10–20 tool-call range, enforced in the runtime. Keep n8n or Dify around for triggers, connectors and the human-facing pieces; rewriting 400 integrations is not a good use of anyone's month. Temporal the moment a single agent run exceeds a few minutes.
What I would not use: a drag-and-drop canvas as the source of truth for an autonomous agent. The tools are not weak — n8n's AI nodes are good and Dify's app model is well thought out. But an agent that runs unattended needs a diff, a test and a version, and a JSON graph gives you none of the three in a form a human can review at 2am.
And I would not self-host inference in the first month. AI doing the work by default is a claim about who performs the task, not about where the weights sit. Running your own GPUs is a cost optimisation with a compliance side effect, and optimisations you make before you have the meter reading are guesses with a monthly invoice attached.
Self-hosting does not reduce the number of things that can break. It moves them onto your pager, in exchange for a boundary you can point at in a contract.
Sources
- n8n repository — Primary source for n8n's license file and codebase
- n8n queue mode docs — Scaling n8n past a single process requires Postgres plus Redis and separate worker processes
- n8n advanced AI docs — n8n's native AI Agent and LLM chain nodes
- Dify LICENSE — Dify is Apache 2.0 with additional conditions covering multi-tenant service use and frontend branding
- Dify documentation — Dify apps export to a YAML DSL file
- Langflow repository — Langflow's MIT license and flow-as-JSON export
- Flowise repository — Flowise license and self-hosting instructions
- CrewAI repository — CrewAI is MIT licensed
- LangGraph repository — The LangGraph library itself is MIT licensed
- LangChain pricing — Self-hosted LangGraph Platform deployment sits on the Enterprise plan, unlike the MIT library
- LangGraph recursion limit error — LangGraph enforces a default recursion_limit of 25 super-steps outside the model
- Temporal repository — Temporal server is MIT licensed and self-hostable for durable execution
- Open WebUI repository — Open WebUI's license is a revised BSD-3-Clause with a branding-retention condition
- vLLM documentation — vLLM's Apache 2.0 self-hosted inference server, continuous batching and PagedAttention KV cache management
- Ollama repository — MIT-licensed local model runner used for single-node self-hosted inference
- Llama 3.3 70B Instruct model card — 70 billion parameters and a 128k context window — the basis for the VRAM arithmetic
- NVIDIA H100 product page — 80 GB HBM per H100 SXM card
- OpenAI API data controls — API inputs and outputs are not used for training by default; retention window and zero-data-retention options
- OpenAI enterprise privacy — Retention and ZDR commitments for API traffic
- Model Context Protocol — Open protocol for exposing tools to agents; lets the tool layer be hosted separately from the orchestrator
- LiteLLM repository — MIT-licensed self-hosted proxy that gives one OpenAI-compatible endpoint across providers, with per-key spend tracking