What to automate in AI account-based marketing, and what to gate
Agents are good at account research and scoring, and dangerous on send and spend. Where the write gate belongs, what a per-account research run costs, and what I would not ship.
What the term actually covers
AI account-based marketing is ABM where an agent does the per-account reading, scoring and drafting that no human could do across 2,000 accounts, while a human keeps send authority and spend authority. That is the definition that survives contact with production. Everything else sold under the label is enrichment with a new price tag, or an unsupervised sender that will eventually cost you a domain.
ABM software is an established commercial category: vendors have licensed account selection, audience targeting and orchestration as products for well over a decade, which is long enough for the operational failure patterns to be documented rather than theoretical. Metadata.io went from $0 to $15M ARR, raised $50M, and I hold six patents in AI-driven marketing. I am now building a zero-human company — agents doing the work, me paying the bills. So the split below is not a philosophy. It is where the failures are.
ABM breaks into five jobs. They do not automate equally.
| Job | Agent-ready? | What decides it |
|---|---|---|
| Account selection and scoring | Yes, with a human-reviewed score distribution | Read-only. Errors are recoverable before anyone is contacted. |
| Per-account research (10-K language, hiring pages, product changelogs, org changes) | Yes | Pure read work at a volume humans never attempted. |
| Message and offer drafting | Yes, as draft | Output is text in a queue, not an action in the world. |
| Channel execution — ad spend, email send, LinkedIn DMs | No | Irreversible. Money leaves, mail lands, reputation moves. |
| Measurement and attribution | Partly | Agents can assemble the report. They should not pick the denominator. |
The interesting row is research. Classic ABM was gated on analyst hours: one person could do a real dossier on maybe five accounts a week. An agent that fetches 30 pages, extracts claims with citations and writes a one-paragraph "why now" does that for 2,000 accounts. That is the actual change. Not personalisation at scale — research at scale, which is what personalisation was always waiting on.
Where the write gate goes
Put every irreversible action behind one gate, and put the gate outside the model. Not in a prompt. Not in a system message that says "always ask before sending." In code, as a separate service that the agent cannot call without a human-approved token.
Two things on the left side are load-bearing.
The step cap. Unbounded loops are the most common production failure in agent systems. The fix is a counter enforced by the runtime, not by instruction. LangGraph ships a default recursion limit of 25 super-steps and raises a GraphRecursionError when it is hit (LangGraph docs). The OpenAI Agents SDK exposes max_turns on the runner, plus input and output guardrails that run as separate checks (Agents SDK docs). Use the framework's limit. A research agent that cannot find a company's funding history will try the same four search queries until something stops it, and the thing that stops it should not be your monthly bill. I wrote up the pattern in build a three-agent system that stops when it should.
Evidence attached to every claim. The draft that reaches the gate carries the URL behind each assertion. Without it, a reviewer has to re-do the research to approve the message — which costs more than writing it from scratch, and collapses the whole system back into a human workflow with extra steps.
What a per-account research run costs
This is a scenario, not a measurement from my own logs. A research agent makes roughly 12 tool-using turns per account, carrying about 20,000 input tokens per turn as fetched page content and an accumulated dossier, and emitting about 1,200 output tokens per turn.
| Input tokens | Output tokens | |
|---|---|---|
| Per turn | 20,000 | 1,200 |
| Per account, 12 turns | 240,000 | 14,400 |
| One full refresh, 2,000 accounts | 480,000,000 | 28,800,000 |
Current per-token rates are on the OpenAI pricing page. Move the sliders and pick a model to see what your own turn count does to the total.
Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.
Three things move that number more than model choice does.
Prompt caching moves it most. The dossier prefix is stable across turns within a single account run, and Anthropic bills cache reads at a fraction of base input price with a default five-minute cache lifetime (prompt caching docs). Re-send the growing dossier uncached every turn and you pay full input price twelve times for the same text. Put the stable part first.
Turn count moves it second. Most of the waste described in agent postmortems is retries against a tool that returned an error the model did not understand, not reasoning depth.
The third lever is the meter itself. Vendor platforms bill in credits, not tokens — Clay prices enrichment and research by credit, Apollo by seat plus credit. Credits are fine. What you lose is the ability to see that one account consumed forty calls because its site blocked your fetcher. If you cannot see per-account cost, you cannot kill the accounts that are not worth researching. I would build the meter first, every time, which is the argument in how to make an AI startup, starting with the meter.
The approach I would not use
I would not give an agent send authority on cold email. Not because the writing is bad — the writing is often better than a junior SDR's. Because the action is irreversible and the liability is per-message. CAN-SPAM requires that opt-outs be honoured within 10 business days, and each non-compliant message is a separate civil penalty (FTC compliance guide). An agent that misreads a suppression list and sends 4,000 messages has created 4,000 violations before anyone opens a dashboard. Delivered mail has no rollback. A draft queue costs a reviewer an hour a day and removes the entire class of failure.
I would not let an agent allocate ad spend without a cap enforced at the platform. Not a cap in the agent's instructions. A lifetime budget on the campaign object itself. Spend authority is the same irreversible-action problem with a faster clock.
I would not build account research on scraped LinkedIn. The LinkedIn User Agreement prohibits scraping and unauthorised automated access, and the vendors who do it on your behalf are a dependency that can disappear in a product cycle. Company sites, filings, job boards, changelogs, podcasts and press releases are public, fetchable, and stay fetchable.
I would not buy "AI SDR" as a replacement for a sales team. The research and the first draft transfer. The qualifying conversation does not, yet. I went through what actually transfers in replacing a sales team with AI.
One legal boundary before you ship contact-level scoring in Europe. GDPR Article 22 gives people the right not to be subject to decisions based solely on automated processing where those decisions produce legal effects or similarly significantly affect them (Article 22). Account-level fit scoring is a long way from that. Automated individual decisions about named people, with no human in the loop, are closer than teams assume. Regulation (EU) 2024/1689 adds transparency duties for systems that interact with people directly (EUR-Lex).
Failure modes you should design for before you see them
None of these are exotic.
Unbounded loops. Covered above. Cap enforced by the runner, not the prompt.
Hallucinated account facts entering the CRM as truth. An agent writes "expanding into EMEA" into a custom field. Six weeks later a rep quotes it on a call and there is no way to find out where it came from. The fix is structural: every agent-written field gets a sibling field holding the source URL and the run ID, and any field with an empty provenance sibling is treated as unverified. Cheap to build. Impossible to retrofit once 2,000 accounts are polluted.
Tool errors read as empty results. A 403 from a company's site, returned to the model as an empty string, becomes "no hiring activity found" in the dossier, which becomes a low intent score. Distinguish no data from fetch failed at the tool boundary. The Model Context Protocol spec gives you a consistent place to put that distinction when you are wiring CRM and warehouse access, instead of one bespoke adapter per system.
Audience minimums colliding with small target lists. LinkedIn enforces a 300-member minimum on Matched Audiences company lists, so a 40-account tier-one list cannot be served as its own segment (Matched Audiences). Agents happily produce 40-account segments. The platform refuses them. Tiering is a constraint of the channel, not a strategic preference, and your scoring output should be shaped to it.
Identity drift. The agent researches acme.com, the CRM holds acmecorp.io from a 2021 import, and both get scored. Resolve on domain plus a canonical company ID before research runs, not after. Duplicate research is duplicate spend.
The thing to score the system on
Pick the metric before you build, because the metric decides what the agent optimises toward. Meetings booked is the wrong one. An agent optimising for meetings booked will find the accounts most willing to take a meeting, which correlates with being unemployed, bored, or a competitor. Pipeline created is better and still gameable. Closed revenue per researched account, measured on a cohort with a lag the length of your sales cycle, is the one that cannot be faked by a text generator.
That number tells you the only thing you need to know about your research agent: whether the dossier changed an outcome, or just made the outreach longer. At Silver Spotfire I took a product from $50K to $1.5M ARR without any of this, by talking to people. The agents earn their token bill when the per-account dossier beats that — and you will only know because you measured the cohort, not because the drafts read well.
Sources
- OpenAI API pricing — Per-token input and output prices used in the per-account cost arithmetic.
- Anthropic prompt caching documentation — Cache reads are billed at a fraction of base input price; default cache lifetime is 5 minutes — the mechanism that makes repeated-dossier research runs affordable.
- OpenAI Agents SDK documentation — max_turns and input/output guardrails as limits enforced by the runner, outside the model.
- LangGraph: GRAPH_RECURSION_LIMIT — Documented default recursion limit of 25 super-steps and how to raise it.
- Model Context Protocol — Open standard for connecting models to tools and data sources, including CRM and warehouse servers.
- GDPR Article 22 — automated individual decision-making — Right not to be subject to decisions based solely on automated processing; constrains contact-level scoring in the EU.
- FTC CAN-SPAM Act Compliance Guide for Business — Opt-out must be honoured within 10 business days; each violating email is a separate civil penalty.
- LinkedIn User Agreement — Prohibition on scraping and unauthorised automated access to the platform.
- LinkedIn Matched Audiences — Company-list targeting and the minimum audience size that forces account tiering.
- Regulation (EU) 2024/1689 (AI Act) — Transparency obligations for AI systems interacting with people.
- Clay pricing — Credit-based metering for enrichment and research runs — an alternative meter to per-token billing.
- Apollo pricing — Seat-plus-credit pricing for contact data and sequencing.
- Metadata.io — The company referenced in the author's background.