Replacing a sales team with AI: what actually transfers
A task-by-task judgement on which sales work AI agents can own today, which commits you to a human signature, and the drop-in AI SDR I would not buy.
The short answer
τ-bench, the tool-use benchmark published by Sierra, reports that agents built on frontier models solve well under half of realistic multi-step service tasks, and that running the same task eight times in a row collapses their consistency. Sales work has that exact shape. Many steps, many tools, a human on the other side who changes the plan mid-sentence.
You can replace most of a sales team's tasks and you cannot yet replace its authority — any step where a wrong action is cheap and reversible is agent-ready now, and any step that commits the company to a price, a scope, or a signature still needs a named human on it.
That is not a seniority line. An SDR's calendar invite is reversible; a VP's verbal discount is not. Sort by consequence, not by title.
I took Metadata.io from $0 to $15M ARR with a sales team, raised $50M, and hold six patents in AI-driven marketing. I now build toward a zero-human company, where agents do the work by default. Both of those things are true at once.
Task by task: what an agent can own
| Sales task | Agent-ready today | What breaks when it is wrong |
|---|---|---|
| ICP list building from firmographic filters | Yes, unsupervised | Wasted enrichment credits. Delete the list. |
| Enrichment and dedupe against CRM | Yes, unsupervised | Bad field value. Reversible with an audit column. |
| Account research brief before a call | Yes, unsupervised | A rep reads a wrong fact and corrects it live. |
| CRM hygiene: stage, next step, close date | Yes, with write-scope limits | Forecast noise. Recoverable from field history. |
| Call recording → summary → next actions | Yes, unsupervised | Missed action item. Human reviews the transcript. |
| Inbound reply triage and routing | Yes, with an escalation path | Slow response on a real buyer. Expensive but recoverable. |
| Sequence timing, throttling, suppression | Yes, deterministic code not an agent | Over-send. See the deliverability section. |
| First-touch outbound copy | Draft, human approves at the campaign level | An invented fact in front of 5,000 strangers. Not recoverable. |
| Discovery conversation | No | You lose the deal and never learn why. |
| Pricing quote outside the published table | No | You just set a precedent for the next renewal. |
| Discount or non-standard terms approval | No | Direct revenue loss, plus an audit problem. |
| Contract redline | No | Legal exposure with no undo. |
Notice what moved. The entire research-and-hygiene layer transfers cleanly, and that layer is most of the hours in a sales development function. The commitment layer does not transfer at all, and it is most of the value in an account executive.
The boundary: reversibility, not difficulty
The useful question is not "can the model do this?" Models can draft a redline. The question is what happens in the 1-in-20 case, and whether you can get the money back.
Two inputs decide it: how much a wrong action costs, and how hard it is to unwind. Plot a task on those axes and the policy falls out.
The solid line is the one that matters. Below it, a mistake is an annoyance you fix with a database update. Above it, a mistake is a commitment somebody can hold you to. I will not move spend authority, pricing authority, or contract language above that line for the sake of a headcount number.
The dashed line is softer and moves with your eval coverage. Outbound copy sits between the two lines because the content is reviewable and the send is not. Once 5,000 people have read a hallucinated claim about their own company, there is no retraction.
The approach I would not use: the drop-in AI SDR
A sealed "AI SDR" that signs up, connects your inbox, and starts sending. I would not do it. Four reasons, all documented rather than aesthetic.
The liability does not transfer with the work. The FTC's CAN-SPAM guidance is explicit that both the company whose product is promoted and the company that actually sends the message can be held legally responsible, and that opt-outs must be honoured within 10 business days. You are buying volume and keeping the exposure.
Deliverability is your asset, and the vendor is spending it. Google's sender guidelines define bulk senders at 5,000 messages per day to Gmail accounts and tell you to keep spam rates below 0.10% and never above 0.30%. That is a per-domain reputation you spent years building. A tool optimised for sends-per-seat has no stake in it.
Voice is a regulatory wall, not a product gap. On 8 February 2024 the FCC ruled that AI-generated voices in robocalls count as "artificial" under the TCPA, which means prior express consent. Any vendor selling an AI voice SDR that cold-calls a purchased list is selling you a compliance problem with a nice dashboard.
You cannot inspect the loop. No step cap you control, no logs you own, no eval set, no way to diff behaviour when the vendor silently swaps the underlying model. That last one is the quiet killer: a sealed product's quality changes on a Tuesday because someone else changed a default. I wrote about which parts of an agent stack you actually control in self-hosted AI agent builders.
Read the pricing shift as a signal. Salesforce prices Agentforce on consumption rather than per seat, and when vendors stop selling seats, "replace the team" stops being a licence conversation and becomes a run-cost conversation. Which is the right conversation.
What the working architecture looks like
The version that holds up in production is boring: a deterministic pipeline with small agentic pockets, not one agent with a long prompt and a company credit card. Anthropic's own guidance on building effective agents says the same thing in different words. Use the simplest composition that works, and reach for an autonomous loop only where the path genuinely cannot be specified in advance.
- Orchestration in code. Triggers, throttles, suppression lists, send windows, routing rules. These are
ifstatements. Do not pay a model to re-derive them 5,000 times a day. - Agentic pockets for open-ended steps. "Read these six sources and write a 120-word account brief" is a real agent task. "Decide what to do about this account" is not a task, it is a job description.
- Tools, not scraped UIs. Expose CRM reads, enrichment, and calendar through a tool layer — MCP is the standard worth building against — with write scopes narrowed to specific objects and fields.
- A step cap enforced outside the model. The OpenAI Agents SDK takes
max_turnsand raisesMaxTurnsExceeded; LangGraph's executor carries a recursion limit that raises rather than spinning. Unbounded loops are the most common production failure and the most expensive, because a loop burns tokens at full speed with no output. - Guardrails as separate checks. The Agents SDK runs input and output guardrails alongside the agent and can halt it. A model asked to grade its own output is not a control.
- State in the warehouse, not in the context window. Long contexts degrade: accuracy drops when the relevant fact sits in the middle of a big prompt. Retrieve the four fields the step needs. Do not paste the account's entire history and hope.
- An idempotency key per contact per campaign. The cheapest line of code in the system. It is what stops the same prospect getting the same email three times after a retry.
For the fuller argument on what "AI does the work by default" means structurally, see AI-first company means AI does the work by default.
Price the run before you cut the headcount
"Replace the sales team with AI" is an arithmetic claim, and most people making it have never priced a single run.
Here is a realistic shape for one unit of work: an account-research-and-draft agent that pulls from a handful of sources, writes to the CRM, and produces a brief plus a draft first touch. Call it 40 tool-calling turns, roughly 15,000 input tokens per turn once you include the tool schemas and accumulated history, 1,500 output tokens per turn, across 5,000 accounts a month. Published per-million rates from OpenAI and Anthropic are what you multiply.
Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.
Three things to do with that number.
Drag the turn count up. Cost scales with turns and input grows as history accumulates, so a loop that wanders from 40 turns to 120 does not cost 3× — it costs more, because every turn re-sends a longer context. The step cap is a budget control, not a reliability control.
Then look at cached input pricing. Stable prefixes — system prompt, tool schemas, product facts — are the bulk of input tokens, and the cached rate is a fraction of the standard one.
Then add everything the model meter does not see: enrichment and data credits billed separately (Clay prices in credits, not tokens), email infrastructure, your engineering hours building and maintaining evals, and the review time of whoever approves what crosses the line in the diagram. Compare that total against your own fully loaded cost per rep. I am not going to invent your salary numbers; you have them. The meter-first habit is the one I keep coming back to, and I wrote it up in build a SaaS app with AI, starting with the token meter and start an AI company by building the meter first.
The failure modes you inherit
None of these are exotic. All of them show up the week you go from 50 runs a day to 5,000.
Unbounded loops. An agent retries a failing tool, re-plans, retries again. Three independent limits stop it: a hard turn cap in the runner, a wall-clock timeout, and a per-run token budget. Three, because any one of them can be the wrong unit.
Invented specifics in outreach. Models are trained and evaluated in ways that reward a confident guess over an admission of uncertainty — OpenAI's own write-up on hallucination makes that the central explanation. In sales copy it surfaces as a fabricated funding round, headcount, or tech-stack detail. The fix is structural: every factual claim in a draft carries a source field populated by a retrieval step, and a guardrail drops the sentence when that field is empty.
| Failure | Fix |
|---|---|
| Context rot: accuracy falls when the needed fact is buried mid-prompt | Narrow retrieval per step, task instruction last |
| Lost updates: agent reads a record, a rep edits it, agent writes back over the edit | Conditional writes on a version field; never let an agent write a field a human owns |
| Duplicate sends after a retry | Idempotency key, checked before send, not after |
| Silent vendor drift: a default model or schema changes and output quality shifts with no deploy on your side | Pin versions where the API allows it, watch the changelog, run a golden set of 50 real accounts nightly with a diff |
If you cannot run that nightly diff, you do not have an automated function. You have an unmonitored one.
What stays human, and who gets paged
Pricing. Scope. Contract language. Anything with a signature or a wire transfer attached. I will not delegate spend authority to an agent yet, and the reason is the τ-bench result at the top of this piece: the failure rate on multi-step tool work is not 1-in-1000, and repeated attempts do not converge. I went through the specific parts of a GTM motion I would leave alone in agentic go-to-market, and the parts I would not automate.
The part most teams skip is the last one. When the agent sends the wrong quote to the wrong account, a person's name has to be on the incident. Not the vendor's. Not "the AI's." A function with nobody on call for it has not been automated — it has been abandoned, and you find out at the quarter close.
Sources
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — Frontier-model agents solve well under half of realistic tool-using service tasks; reliability across eight repeated trials (pass^k) falls sharply.
- Anthropic — Building effective agents — Distinction between deterministic workflows and agents; guidance to use the simplest structure that works and add agentic loops only for open-ended steps.
- OpenAI Agents SDK — Running agents — max_turns parameter and MaxTurnsExceeded: a step cap enforced by the runner, outside the model.
- OpenAI Agents SDK — Guardrails — Input and output guardrails run alongside the agent and can halt execution, i.e. validation outside the model's own judgement.
- LangGraph repository — Graph executor with a recursion limit that raises instead of looping indefinitely.
- Lost in the Middle: How Language Models Use Long Contexts — Retrieval accuracy degrades when the relevant information sits in the middle of a long context.
- OpenAI — Why language models hallucinate — Training and evaluation reward confident guessing over abstention, which is why models invent specifics.
- FTC — CAN-SPAM Act Compliance Guide for Business — Opt-outs must be honoured within 10 business days; both the company whose product is promoted and the company that sends the message can be held legally responsible.
- FCC — AI-generated voices in robocalls are illegal — Declaratory ruling of 8 February 2024 making AI-generated voices 'artificial' under the TCPA, so such calls require prior express consent.
- Google — Email sender guidelines — Bulk sender threshold of 5,000 messages per day to Gmail accounts; keep spam rates below 0.10% and never above 0.30%.
- Model Context Protocol — Open standard for exposing tools and data sources to models.
- Anthropic pricing — Per-million-token input and output pricing used for run-cost arithmetic.
- OpenAI API pricing — Per-million-token pricing, including cached input rates.
- Salesforce Agentforce pricing — Agentforce is priced on consumption rather than per seat.
- Clay pricing — Enrichment and data credits are billed separately from model tokens.