What Harvard actually measured about AI and entrepreneurship
Two Harvard field experiments — 758 consultants and 776 P&G professionals — give founders a routing rule, not a forecast. Here is the rule, and where it stops working.
The one number that should change how you staff a company
In a field experiment with 758 Boston Consulting Group consultants, the group given GPT-4 completed 12.2% more tasks, moved 25.1% faster, and produced work rated more than 40% higher in quality. On a task deliberately designed to sit just outside the model's capability, the same tool made them 19 percentage points less likely to reach the correct answer than colleagues working without it (HBS Working Paper 24-013).
Harvard's contribution to this subject is a measured boundary, not a forecast. AI sharply raises output inside its capability frontier and degrades it just outside. The frontier is jagged, not a clean line. A founder's actual job is routing work across it.
The paper is Navigating the Jagged Technological Frontier, from a team including Fabrizio Dell'Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon and Karim Lakhani. It is a working paper, free to download. It is also the single most quoted piece of Harvard work on this topic, and most people who cite it quote the upside half and drop the 19 points.
Drop the 19 points and you have a productivity pitch. Keep them and you have an operating constraint.
The five different Harvards in your search results
"Entrepreneurship in the age of AI" plus "Harvard" returns at least five unrelated things. They cost different amounts, admit different people, and hand you different artifacts.
| What it is | What you get | Who can access it |
|---|---|---|
| Digital, Data, and Design Institute research | Working papers and field experiments, including the two below | Anyone, free |
| HBS MBA electives and course catalog | Degree-program courses, case discussions, faculty office hours | Admitted degree students |
| HBS Executive Education | Short paid programs, days to weeks, on campus or live online | Anyone who pays |
| HBS Online | Self-paced certificate courses | Anyone who pays |
| Harvard Innovation Labs | Space, mentorship, venture programs | Harvard student and alumni ventures only |
If you are a founder outside Harvard and you want the substance rather than the credential, the research tier is the only one that matters. It is also the only free one. Executive Education sells you a cohort and a certificate, and neither of those shows up in your product.
Two papers carry almost all the signal. The first is the jagged-frontier study above. The second is The Cybernetic Teammate, a field experiment with 776 professionals at Procter & Gamble working real product-development problems, structured as individuals versus teams, with AI versus without.
Individuals with AI matched teams without AI
That is the headline result of the P&G experiment, and it is the one that should sit in a founder's head when they open a job requisition. In the reported design, a single professional working with AI produced solution quality comparable to a two-person cross-functional team working without it. Teams using AI were the most likely to produce top-decile solutions. AI use also blurred functional boundaries: R&D participants produced more commercially framed solutions, and commercial participants produced more technical ones (arXiv:2503.18238).
The implication is narrower than the LinkedIn version.
What the paper found is that AI substitutes for some of what a second functional perspective provides. That is not the same as substituting for a person. Nobody at P&G handed the model a budget, a vendor relationship, or the authority to ship. The tasks were bounded, scored, and finished inside a session.
I have been on the hiring side of this question for a decade. I took Metadata.io from $0 to $15M ARR, raised $50M, and hold six patents in AI-driven marketing; before that I grew Silver Spotfire from $50K to $1.5M ARR. Now I am building a zero-human company in public — agents do the work, I pay the bills for them. The P&G result argues for a smaller founding team, not a zero-person one. A two-person team with agents covers ground that used to need six. Day one, that is the whole prize, and it is a large one.
What the paper does not license is a headcount plan built on a number from someone else's controlled study.
Where the Harvard numbers stop
Both experiments are session-scoped: a defined task, a scoring rubric, a human present. Production is not session-scoped. Two primary sources mark the edge.
| Source | Design | Result |
|---|---|---|
| METR, July 2025 | Randomised controlled trial, 16 experienced open-source developers, 246 real tasks in their own repositories | 19% slower with early-2025 AI tools. They had forecast a speedup, and afterwards they still believed they had been sped up |
| Anthropic, October 2024 | Computer use shipped with the upgraded Claude 3.5 Sonnet, OSWorld screenshot-only benchmark | 14.9% against 72.36% for humans, next-best system at 7.8% |
The perception gap in the METR row is the part to carry forward. Self-report on AI productivity is unreliable in both directions, which means your team's opinion of how much the tools help is not evidence.
The OSWorld number was best in class at the time and still roughly a fifth of human performance. If your plan routes work through a GUI because there is no API, that is the number you are building on — a computer use agent drives a GUI with pixels, not APIs, and the pixel path is the most brittle path in the stack.
Neither result contradicts the Harvard papers. They measure a different regime: unbounded task horizon, real repository, no rubric, no observer.
The routing rule the research implies
The jagged frontier is not an idea to agree with. It is a dispatch table you write down and maintain.
The boundary is empirical and local. Your tasks, your data, your tolerance for a wrong answer. A task that sits inside the frontier for a consulting deliverable sits outside it for a regulated filing, and the difference is the cost of being wrong once. You find that out by running the same task ten times and counting failures. Benchmark tables will not tell you.
The middle column is the expensive one. Work at the edge produces plausible output that a domain expert has to check, and checking costs as much as doing. Founders underprice this every time. The BCG consultants who did best treated the model as a colleague to argue with rather than an oracle to copy.
And the boundary moves on someone else's release schedule. A task outside the frontier in March is inside it by October, which makes re-measurement a recurring line item rather than a one-off.
What the agent tier costs before you commit to it
Most founders reach the unit-economics conversation a quarter late.
The arithmetic is not hard. Take a research agent that runs a dozen tool-calling turns, carries a context that grows to roughly 20,000 tokens by the later turns, emits about 800 tokens of output per turn, and runs 2,000 times a month. That is 240,000 input tokens and 9,600 output tokens per run, so about 480 million input tokens a month. Multiply by the per-million input price on the OpenAI or Anthropic pricing page for the tier you picked.
Two things dominate the total, and the output price is neither of them. Input tokens scale with turn count, because every turn resends the transcript. Retries are the other: an agent that fails and reruns pays twice for the same result.
Prompt caching is the lever on the first. Anthropic publishes cache reads at 0.1× base input price and cache writes at 1.25×, so a stable system prompt and tool block turn the dominant cost line into a tenth of itself across a long multi-turn loop (Anthropic pricing).
Arithmetic on published list prices. Retries are billed too, so a step that fails twice before working costs three times this.
Move the run count to what you would actually charge for and see whether the margin survives. If it does, that is a product: ship one metered agent endpoint and bill per run with something like Stripe Billing rather than selling seats for work no human performs. If the margin does not survive, you learned that for ten minutes instead of a quarter. More on the general shape of this in what AI can actually do for a business, and what it costs.
The approach I would not use
I would not start by drawing an org chart of agents.
It is the most common design I see from founders who have read the research and want to act on it: a "CMO agent" that delegates to a "content agent" and a "research agent," each with a persona, coordinating through free-text handoffs. It demos beautifully. It fails in a boring, specific way — unbounded loops. Two agents pass work back and forth, neither has a termination condition the model cannot rationalise away, and the loop burns tokens until something external kills it. This is the most frequently documented production failure in agent systems, and the fix is not a better prompt. It is a step cap enforced outside the model, plus a hard budget per run. LangGraph ships a recursion limit for exactly this reason; the OpenAI Agents SDK exposes max-turns and guardrails as first-class primitives. Use them. The minimal version is in three tools, a step cap, a JSONL note, and the definitional question of what even counts as multi-agent in multi-agent system, defined without the hype.
Second, I would not delegate spend authority. Not yet. An agent that can call a paid API, buy ads, or issue a refund needs an approval gate the model cannot route around, because the failure mode is not a bad sentence, it is a charge on a card. The P&G and BCG results say nothing about agents holding a budget. In neither study did they.
Third, I would not pick tools from benchmark leaderboards. OSWorld at 14.9% against a human 72.36% tells you the category is early. It tells you nothing about whether a specific vendor's agent finishes your specific workflow. Run your own task set and count failures.
What I would build instead is narrower than an org chart and duller than a demo. One loop that decides its own next step — that is the actual test of whether something is an agent — with three tools, a step cap, a structured log per run, and one human review gate placed where a wrong answer costs real money. Connect it to data through a documented interface rather than a scraped one: the Model Context Protocol exists so tool access is a declared contract instead of a brittle integration. Add the second loop only after the first has run unattended for a few weeks and you have a failure log to read. That sequencing — one loop proven in production before a second one exists — is what the bounded findings above actually license, since neither experiment measured an unsupervised system running past the end of a session; a worked account of the approach is in running a company with AI agents.
The Harvard papers are useful precisely because they are bounded. They measured a session, scored the output, and published the number that made the tool look bad alongside the three that made it look good. Hold your own systems to that standard and you will discard most of what you build. That is the correct outcome.
Sources
- The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise — 776 Procter & Gamble professionals; individuals with AI matched teams without AI; AI use eroded functional silos
- Digital, Data, and Design Institute at Harvard — The Harvard institute that hosts the research groups behind both field experiments
- Harvard Business School MBA program — Degree-program curriculum and electives, as distinct from executive or online offerings
- Harvard Business School Executive Education — Paid, short-format programs open to non-degree participants
- Harvard Business School Online — Self-paced certificate courses, including AI and analytics subjects
- Harvard Innovation Labs — Venture support restricted to Harvard student and alumni ventures
- METR: Measuring the impact of early-2025 AI on experienced open-source developer productivity — 16 experienced developers, 246 tasks; 19% slower with AI tools while believing they were faster
- Anthropic: Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku — OSWorld scores: 14.9% for the model versus 72.36% for humans
- Anthropic pricing — Per-million token prices and prompt caching multipliers
- OpenAI API pricing — Per-million token prices used for agent run-cost arithmetic
- Model Context Protocol — Open standard for connecting models to tools and data sources
- OpenAI Agents SDK (repository) — Primary source for agent loop, handoff and guardrail primitives
- LangGraph (repository) — Graph-based orchestration with explicit state and recursion limits
- Stripe Billing — Usage-based and metered billing, for pricing an agent endpoint per run
- Navigating the Jagged Technological Frontier (HBS Working Paper 24-013) — 758 BCG consultants; 12.2% more tasks, 25.1% faster, >40% quality lift inside the frontier; 19 percentage points less likely to be correct outside it