An overhead shipyard with four crane-equipped slipways feeding identical vessels along a magenta route through an inspection gate into open water.A software factory builds agent work in parallel and sends it through inspection before release.

Every founder we talk to asks a version of the same question. "My engineers already use AI. Why isn't the company shipping twice as fast?"

The short answer: an AI coding assistant makes a person faster at writing code. A software factory makes a company faster at shipping working software. Those are different things. The research from the last two years explains the difference, and it makes the business case for building the second one.

This post is for CEOs, CTOs, product founders and solo builders. It covers what a software factory is, what the data says about the value, where the value leaks away, and how to start at your stage. Every number is sourced at the end. Where the evidence cuts against the story, we show that too.

The one-paragraph version. AI agents can now do multi-hour engineering tasks, and that capability is doubling roughly every four months. Adoption is close to universal. But the gains show up in company results only when the work around the code is redesigned: queues, verification, review, merge, budgets and memory. Without that, AI produces more code, more review load and more defects, and the business barely moves. A software factory is the system that turns agent capability into shipped, trustworthy product.

1. What a software factory is

A software factory is a governed production line for software, where AI coding agents do much of the building and people decide what gets built and what ships.

Eight steps connect goals and ideas to production, with policy gates, verification and risk-based review. Shared memory supports every step; measured outcomes feed back to goals.What a software factory is

It has six parts, and each one exists because agents fail in a specific way without it:

Part What it does The failure it prevents
Intake and queue Turns goals into small, well-defined work items with owners and priority Agents working on the wrong thing, or two agents on the same thing
Agents Claude Code, Codex and similar tools doing the implementation, each in its own sandboxed workspace One agent's mess spilling into another's work
Policy gates and budgets Deterministic rules on what an agent may run and spend Force-pushes, deleted databases, surprise cloud bills
Verification Tests and checks that must pass before an agent may say "done" "The tests pass" when they never ran
Review and merge Risk-based human review and an exact, auditable merge pipeline Unreviewed code reaching production
Memory What was decided, tried and learned, searchable by every agent Every session re-deriving yesterday's work

The key idea: the agent is not the product. The system around the agent is.

Comparison of human-only, copilot and software factory operating models across code authorship, work units, bottlenecks, governance, scaling and measured gains.Three ways to build software with AI

2. The capability curve: why this is a now problem

In 2023 the best AI model could reliably finish a software task that takes a human expert about 4 minutes. By February 2026, the best models could finish tasks that take an expert about 12 hours, half the time. That is a ~180× increase in three years.

Linear-scale task horizons by model release date: GPT-2 0.05 min; GPT-3.5-instruct 0.6 min; GPT-4 4.0 min; GPT-4o 7.0 min; Claude 3.5 Sonnet 20.5 min; o1 38.8 min; Claude 3.7 Sonnet 60.4 min; o3 119.7 min; GPT-5 203 min; Gemini 3 Pro 224.3 min; Claude Opus 4.5 293 min; GPT-5.2 352.2 min; Claude Opus 4.6 718.8 min. Illustrative least-squares exponential fit to the shown 2023–2026 observations on a linear hours axis; observations remain unchanged.The 50% success horizon rose from GPT-4’s 4 minutes to Claude Opus 4.6’s 718.8 minutes (~12 hours). Linear hours reveal the exponential shape; the curve is an illustrative least-squares exponential fit to the displayed 2023–2026 observations, not a forecast. METR reports ~129-day doubling for its broader trend..")

METR, an independent research organisation, measures this "time horizon": the length of task, in expert-human time, that an AI agent completes with 50% success. Since 2023 it has doubled about every 129 days. Over the full history since 2019, the doubling time is about 188 days.

What this means for a business:

  • A 4-minute task is autocomplete. A person stays in the loop for every line.
  • A 1-hour task is a delegated ticket. Someone reviews the result.
  • A 12-hour task is a day of engineering. You need a queue, a budget, checkpoints and a way to verify the outcome without redoing it.

The tools crossed from the first category to the third in about 30 months. Most companies' processes did not.

A note on benchmarks: public coding benchmarks like SWE-bench Verified went from about 50% to over 90% in under two years, and OpenAI stopped reporting the benchmark in February 2026 after an audit found many of its tasks flawed. Treat any single leaderboard number with caution. The direction is not in doubt; the exact score is.

3. Adoption is near-universal. Trust is not.

Use or plan to use AI: 76% in 2024 to 84% in 2025; distrust of AI accuracy: 31% to 46%; AI agent use: 31% in 2025 to 59% in the April 2026 pulse.AI adoption and distrust both increased, while reported agent use rose to 59% in the April 2026 pulse.

  • 90% of software professionals now use AI at work, a median of two hours a day (Google DORA 2025, about 5,000 respondents).
  • 84% of developers use or plan to use AI tools, up from 76% the year before (Stack Overflow 2025, 49,000+ responses).
  • Use of AI agents, not just chat or autocomplete, rose from 31% in 2025 to 59% in April 2026 (Stack Overflow pulse survey).
  • Google says 75% of its new code is now AI-generated and approved by engineers, up from about 25% in late 2024.

Google 75% (Apr 2026); Robinhood ~50% (Jul 2025); Coinbase ~40% (2025); Microsoft 20–30% (Apr 2025); Uber ~10% by autonomous agents (May 2026). Google trend: 25% Oct 2024, 50% fall 2025, 75% Apr 2026. Executive statements; definitions differ.Reported AI-generated code shares range from ~10% by autonomous agents at Uber to 75% at Google; definitions differ.

The open-source world shows the same shift. GitHub's 2025 Octoverse counts 180 million developers, 1.1 million public repositories using an LLM SDK (up 178% in a year) and 518.7 million merged pull requests (up 29%). The most-starred open-source coding agents were almost all created in 2025: four of them passed 100,000 stars each within about eighteen months.

GitHub stars by repo creation date: opencode: 212,182, Apr 2025; Claude Code: 149,734, Feb 2025; OpenAI Codex CLI: 128,169, Apr 2025; Gemini CLI: 107,245, Apr 2025; OpenHands: 90,185, Mar 2024; MetaGPT: 70,766, Jun 2023; Cline: 69,981, Jul 2024; Goose: 55,037, Aug 2024; Aider: 49,409, May 2023; Continue: 36,148, May 2023; ChatDev: 34,458, Aug 2023; Kilo Code: 27,520, Mar 2025; Roo Code: 24,284, Oct 2024; SWE-agent: 20,497, Apr 2024. Four projects created in 2025 have 597K+ stars combined.Four coding-agent projects created in 2025 passed 100,000 GitHub stars each, reaching 597K+ combined.

At the same time, trust is falling:

  • Only 24% of DORA respondents trust AI output "a lot" or "a great deal".
  • 46% of Stack Overflow respondents distrust the accuracy of AI tools, up from 31%.
  • 66% say the biggest frustration is code that is "almost right, but not quite".

Trust in AI output: 24% a lot / a great deal, 46% somewhat, 30% a little / not at all. 90% use AI; median usage is 2 hours per day.Only 24% of DORA respondents expressed high trust in AI output, despite 90% using AI.

This gap, near-universal use with low trust, is the whole business problem in one picture. People are generating more code than ever and believe less of it. Somebody has to check it, and that somebody is your most expensive engineers.

4. What the evidence says about productivity

Here are the controlled studies and large field measurements, side by side.

Effects: Copilot +55.8% faster; Google +21% faster (~96 vs 114 min); Microsoft/Accenture +26.1% tasks ±10.3; Microsoft CLI +24.0% merged PRs, CI +14.5 to +33.7, daily users +50.1%; Faros +21% tasks; METR −19% slower. METR 2026 follow-up faster on average but inconclusive due to selection bias. METR experts predicted +24%, measured −19%, believed +20% afterwards.Reported effects vary across settings, with METR’s 2025 study finding a 19% slowdown among experienced developers.

Study Setting Result
GitHub Copilot RCT (2023) One scoped task: build an HTTP server 55.8% faster
Google internal RCT (2024) 96 Google engineers, enterprise task ~21% less time
Microsoft, Accenture, Fortune-100 field experiments (2025) 4,867 developers, real work +26% completed tasks; juniors gained most
Microsoft agentic CLI rollout (2026) Command-line coding agents, 4 months +24% merged PRs, no decay; daily users +50%
Faros telemetry (2025) 10,000+ developers, 1,255 teams +21% tasks, +98% PRs, but flat company-level metrics
METR RCT (2025) 16 expert open-source maintainers on their own repos 19% slower, while believing they were 20% faster
METR follow-up (2026) 57 developers, 800+ tasks Faster on average, but too biased by selection to be conclusive

Three things stand out for a business owner.

First, the gains are real and they repeat. Well-defined work, measured inside a real organisation, consistently shows 20–30% more output, and agentic tools in 2026 show similar or better, sustained over months.

Second, people misjudge their own speed. In METR's 2025 study, experienced developers predicted AI would make them 24% faster, were actually 19% slower, and still believed afterwards they had been 20% faster. Self-reported productivity is not a metric. If you don't measure, you don't know.

Third, individual output and company output are different numbers. Faros saw teams merge almost twice as many pull requests, and saw no improvement at the company level. McKinsey's 2026 Technology Trends Outlook reports the same pattern across companies: AI tools raised coding activity by 180%, but shipped releases rose by only 30%; only a quarter of companies achieved meaningful acceleration, and nearly 30% saw productivity fall after adopting agentic tools. The next section explains why.

5. What it looks like when it works

The companies getting real results did not just hand out licences. They built pipelines.

Airbnb: 1.5 years estimated to 6 weeks, 97% automated, ~3.5K files. Amazon: ~50 developer-days per app to hours, 30,000 apps, ~4,500 developer-years and ~$260M/yr saved. Google: 6× faster migrations. Stripe: 1,300+ fully agent-written PRs merged weekly, all human-reviewed. Uber: background agent opens 11% of PRs; AI review covers 90%+ of ~65K weekly diffs.Company reports describe faster migrations and substantial agent output, alongside review practices.

  • Stripe merges more than 1,300 pull requests a week written end-to-end by its internal agents, every one reviewed by a person. Work starts as a request in Slack and ends as a PR that has already passed CI, built on isolated environments and hundreds of internal tools.
  • Airbnb migrated about 3,500 test files to a new framework in 6 weeks, against an original estimate of 1.5 years, with 97% of the work automated.
  • Amazon used its coding agent to upgrade 30,000 applications to a new Java version, cutting each upgrade from about 50 developer-days to hours, and estimates 4,500 developer-years and $260 million a year saved.
  • Google reports a large code migration done by agents and engineers together six times faster than a comparable one done by people alone.
  • Uber's background agent opens 11% of all its pull requests, and its AI reviewer covers more than 90% of about 65,000 changes a week.

The flip side is at Uber too: the company spent its entire 2026 AI-coding budget in about four months, and its COO said publicly that the link between agent usage and shipped features "is not there yet". Same tools, same company: where there is a pipeline, there are results; where there is only usage, there is a bill.

6. Where the value leaks: the review bottleneck

When agents write more code, every downstream step gets more work: review, testing, merging, deploying, fixing.

High-AI teams vs baseline: tasks completed +21%, PRs merged +98%, PR size +154%, review time +91%, bugs per developer +9%. Company-level delivery metrics: flat.High-AI teams produced larger PRs and spent more time on review while company-level delivery metrics stayed flat.

In Faros's data, high-AI teams merged 98% more PRs, but their PRs were 154% larger, review time went up 91%, and bugs per developer rose 9%. The constraint moved from writing code to verifying it, and verification capacity didn't grow.

It gets worse over time. A 2026 study of 400 reviewers and 11,429 reviews found that reviewers approve agent-written PRs more and more easily: approval rose from 27.9% to 42.4% across reviewers' own history, while their inline comments fell 22%. Over the same period they became stricter with human PRs.

Agent PR approval rates by reviewer experience decile 1–10: 27.9%, 27.1%, 30.0%, 28.1%, 33.3%, 35.6%, 32.6%, 34.2%, 38.6%, 42.4%. Approvals +14.5 points; inline comments −22%.Approval rates rose from 27.9% to 42.4% across reviewer experience deciles as inline commenting declined.

Dozens of small boats converge across a wide bay toward a narrow harbour mouth guarded by one inspection tower, with only a few boats beyond it.Output grew; verification capacity did not.

On public GitHub, the review gap is measurable. Coding agents opened millions of pull requests in the past year (Codex about 5.1 million, Claude Code about 4.2 million, Copilot about 1.9 million). Most of them merge, but for two of the three biggest agents fewer than a third of merged PRs got a human review. And merging isn't the end of the job: merged agent PRs need follow-up fixes at 1.62× the odds of human PRs, and a Stanford study of 6,000 real agent sessions found that only 44% of agent-written code survives into commits.

PRs opened: Codex 5.1M; Claude Code 4.2M; Copilot 1.9M; Jules 1.2M; Cursor 614K; Devin 161K. Merge rate / human review among merged PRs: Claude Code 91.5% / 30.5%; Codex 87.5% / 28.6%; Copilot 68.1% / 80.5%. Merged agent PRs need follow-up fixes at 1.62× the odds of human PRs.Most merged agent PRs were never reviewed by a person; merged agent PRs had 1.62× the odds of follow-up fixes.

Developers feel both sides of this. Atlassian's 2025 developer experience survey found 68% of developers save more than 10 hours a week with AI, while 50% lose more than 10 hours a week to organisational friction. The time AI gives back is easily eaten by the system around it.

And the quality signals are mixed:

  • Veracode found security flaws in 45% of AI-generated code tasks (over 70% in Java).
  • Apiiro saw 3–4× more code produce 10× more security findings.
  • GitClear measured an 8× rise in duplicated code blocks and a 40% drop in refactoring.
  • On the other hand, a study of 37,623 public agent PRs found security-smell rates lower than humans' when pooled, with large differences between agent vendors.

Veracode: 45% of AI code tasks introduced a vulnerability, Java >70%. Apiiro: 10× more security findings for 3–4× more code. GitClear: 8× more duplicated code blocks, −40% refactoring. A 37,623-PR study behind good gates found agent security-smell rates below humans.Studies report security and maintainability risks, while a gated PR study found lower agent security-smell rates than humans.

The lesson isn't "AI code is bad". It is: the outcome depends on the system the code goes through. The same agents, behind good gates, produce good results. Behind no gates, they produce volume.

7. AI is an amplifier

Google's DORA research summarises two years of data in one sentence:

"AI's primary role is that of an amplifier. It magnifies the strengths of high-performing organisations and the dysfunctions of struggling ones."

Teams with clear policies, small batches, solid version control, good internal platforms and accessible data get the gains. Teams without them get more instability. DORA reports that throughput is up across the industry while delivery instability keeps rising.

McKinsey reaches the same conclusion from the other side: top performers see 16–30% improvements in team productivity, customer experience and time to market, and 31–45% in software quality. What separates them is how they work, not which tools they bought.

AI coding agents amplify strong foundations into faster, safer delivery, and weak foundations into more code and more instability.AI is an amplifier

This is why a software factory is a management decision, not a tooling decision. The factory is how you make your organisation the kind that AI amplifies upward.

8. Where the business value comes from

Five branches of business value: throughput, lead time, quality and risk, cost control, and talent leverage, with the figures supplied in the brief.Where the business value comes from

For a CEO, the value of a software factory shows up in five places:

  1. Throughput. More finished, merged work per engineer: +20–30% in controlled studies, +24% for agentic tools at Microsoft, with Gartner expecting agent-based workflows to reach 30–50% team gains versus 0–20% for assistants alone.
  2. Lead time. Shorter time from idea to production, because the queue, verification and merge steps are automated and run around the clock. Google reports a large migration running "six times faster" with AI.
  3. Quality and risk. Verification that can't be skipped, deterministic policy on dangerous actions, and an audit trail. This is what keeps the 45%-flawed-code statistic out of your production systems.
  4. Cost control. Agent spend is real, and it is growing even as models get cheaper (see the next section). Budgets, model routing and avoiding wasted agent work are part of the factory. One 2026 study found simple developer-written "skills" cut agent cost by 8–42% in most settings.
  5. Leverage of scarce talent. Your senior engineers stop typing and start specifying, reviewing and designing. The factory multiplies the people you can't hire more of.

Cursor annualised revenue run-rate: $2B February 2026, $3B late April 2026, $4B+ June 2026; Claude Code >$2.5B early 2026. Gartner enterprise AI coding agent market ≈ $10–11B/year in 2026; 90% of enterprise engineers to use AI code assistants by 2028, from <14% early 2024.Reported coding-agent run-rates reached billions of dollars as Gartner forecast widespread enterprise adoption.

The market is moving accordingly. Gartner published its first Magic Quadrant for Enterprise AI Coding Agents in May 2026 and sizes the market at roughly $10–11 billion a year. Claude Code alone passed a $2.5 billion run-rate in early 2026, and Cursor reported growth from $2 billion to over $4 billion run-rate between February and June 2026. Enterprises are paying for this, at scale.

9. The economics: cheaper tokens, bigger bills

First, a paradox every CFO will meet: tokens get cheaper, bills get bigger. The price of a fixed level of AI capability falls roughly 10× a year (a16z) and by a median of 50× a year on Epoch AI's measures. Yet spend per developer is rising: Anthropic's own estimate of a typical Claude Code day went from about $6 to about $13 in a year, a quarter of technology leaders already spend $200–500 per developer per month (Gartner), and Gartner expects AI coding cost per engineer to exceed the average developer's salary by 2028. Cheaper intelligence means agents do longer, more autonomous work, and that work costs money. Gartner's warning is blunt: without a governed engineering operating model, costs can escalate faster than the productivity gains.

Constant-capability cost per million tokens: $60 in 2021 to $0.06 in 2024, ~10× cheaper yearly (a16z); Epoch AI median 50× yearly. Claude Code typical cost per developer per active day: $6 in 2025 to $13 in 2026. A quarter of tech leaders spend $200–500 per developer per month. Uber exhausted its 2026 AI-coding budget in ~4 months. Gartner forecasts AI coding cost per engineer exceeding average developer salary by 2028.Lower token prices coexist with rising bills. Budgets belong inside the factory.

We won't give you a made-up ROI number. Here is the model we use with clients instead, so you can plug in your own figures.

Engineer cost times measured throughput gain yields extra-output value. Subtract agent spend, review time and factory operation, then adjust for quality to find net quarterly team value.A simple economic model (use your own numbers)

  • Start from the fully loaded cost of an engineer. The US median software developer wage alone is about $133,000 (BLS, 2024), before benefits and overhead.
  • Apply a measured, conservative throughput gain. 20–25% is what the controlled studies support for well-run teams.
  • Subtract the new costs. Agent and model spend, more review time, and the time to build and run the factory itself.
  • Adjust for quality. Escaped defects and security incidents cost far more than the code that caused them, so verification is worth paying for.

Run the numbers per team, measure for a quarter, and decide with data. The METR study is the reason: if you rely on how productive people feel, you can be off by 40 points.

10. What this means at your stage

The software factory looks different for a solo founder and for an enterprise. The principles are the same.

Four growth stages compare constraints, initial factory capabilities and payoffs, with governance increasing from solo founder to enterprise.The software factory at every stage

Solo founders and indie builders

Solo founding is rising fast: the share of new startups with a single founder went from 23.7% in 2019 to 36.3% in the first half of 2025 (Carta). In Y Combinator's Winter 2025 batch, a quarter of companies had codebases that were about 95% AI-written.

Solo founders accounted for 23.7% of new startups in 2019 and 36.3% in H1 2025. 25% of YC Winter 2025 had codebases ~95% AI-generated. Solo founders received 14.7% of priced-round cash in 2024.The solo-founder share of new startups rose to 36.3% in H1 2025, while YC reported extensive AI-generated code in a quarter of its winter batch.

The economics of small teams have changed. Base44, a bootstrapped company with one founder, was sold to Wix for $80 million in cash six months after launch. Lovable reported about $2.7 million of revenue per employee in early 2026, against a median of about $141,000 for private SaaS companies. Gartner expects 60% of organisations to run "tiny" software teams of two to five people by 2029, up from 15% in 2026, and warns that cutting junior roles weakens the talent pipeline.

Revenue per employee: Lovable ~$2.7M (Feb 2026, company statement) versus median private SaaS $141K (2026, SaaS Capital). Base44: solo-founded, sold to Wix for $80M cash after 6 months (Jun 2025). Gartner forecast: organizations running tiny 2–5 person software teams rise from 15% in 2026 to 60% in 2029.Lovable reports ~$2.7M revenue per employee; Base44 sold for $80M after six months; Gartner forecasts more tiny software teams.

A single faceless founder at a helm beside a lighthouse directs identical boats along coordinated magenta courses across a nautical chart.One founder sets the course for a fleet of agents.

For a solo founder, the factory is a force multiplier: you are the product owner, the reviewer and the release manager, and the agents are your team. What you need most:

  • A queue and small tasks, so agents work on the right thing in the right order.
  • Automatic verification, because you can't review everything by hand.
  • Budgets, because a runaway agent loop is your burn rate.
  • Memory, so you don't re-explain your product to every new session.

Startups and scale-ups

Your constraint is senior engineering time. The factory turns your best engineers into specifiers and reviewers and lets agents absorb implementation, migrations, tests and maintenance. Watch the review bottleneck most closely: if PRs pile up, add risk tiers (low-risk changes need one review, infrastructure needs more) rather than more reviewers.

Enterprises

Your constraint is risk and governance. You need policy as code, audit trails, approvals routed to the right people, spend caps per team, and evidence that controls work. The Gartner Magic Quadrant exists because this buyer exists. The factory is how you let hundreds of agents work without losing control of what reaches production.

11. How to start: a 90-day plan

A three-phase roadmap: measure and pilot in days 1–30, add factory parts in days 31–60, and scale what the data supports in days 61–90, with data-driven checkpoints.Your first 90 days

Days 1–30: measure and pilot. Pick one team and one workflow, such as test writing, dependency upgrades or internal tools. Baseline the metrics you care about: lead time, merged work, escaped defects, review time and spend. Give agents sandboxed workspaces and a short list of things they may never do.

Days 31–60: add the factory parts. Put work into a queue with small, well-defined items. Make verification mandatory: an agent's turn isn't finished until the declared checks pass. Introduce risk-tiered review and a merge pipeline that only merges the exact commit that was verified and reviewed. Set budgets.

Days 61–90: scale what the data supports. Compare against the baseline. Expand to more teams only where the measured numbers moved. Add memory so agents stop repeating work. Report value in business terms: lead time, quality, cost per change.

Boats follow two orderly magenta lanes between navy buoys through a harbour gatehouse with a green signal and an open logbook on the pier.Governance provides clear lanes, permission signals, and a record of passage without slowing traffic.

12. The honest risks

A software factory is not magic, and the research is clear about what can go wrong:

  • Perception is unreliable. Measure outcomes, not sentiment.
  • Review quality decays. Reviewers get used to approving agent work. Make verification automatic and keep human review focused where risk is high.
  • Security and maintainability. AI code can multiply vulnerabilities and duplication without the right gates.
  • Understanding. Developers who let agents do everything understand their code less. Keep people in the design and review loop.
  • Cost. Without budgets and routing, agent spend can grow faster than the value.

Every one of these risks is a reason to build the factory, not a reason to avoid AI. The companies that will lose are the ones that adopt agents without the system around them.

13. How f200 approaches this

We build and run AI software factories, for ourselves and for clients. Our own engineering has run on Shiftyard, our factory for coding agents, since October 2026. It is pre-release and will be open source under Apache-2.0; you can read the technical deep-dive. For clients, our Software Factory Integration service designs and installs the queue, gates, verification, review and memory around the agents your teams already use.

If you want to know what a software factory would mean for your company, at your stage, book a 30-minute call with the founders. We'll start with your workflow and your numbers, not with a tool.

Sources

  1. Peng et al., The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv 2302.06590, Feb 2023. https://arxiv.org/abs/2302.06590
  2. Paradis et al., How much does AI impact development speed? An enterprise-based randomized controlled trial, arXiv 2410.12944, Oct 2024. https://arxiv.org/abs/2410.12944
  3. Cui, Demirer, Jaffe, Musolff, Peng, Salz, The Effects of Generative AI on High-Skilled Work, Management Science, 2025. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4945566
  4. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, Jul 2025; and uplift update, Feb 2026. https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf · https://metr.org/blog/2026-02-24-uplift-update/
  5. Murphy-Hill, Butler, Savelieva (Microsoft), Adoption and Impact of Command-Line AI Coding Agents, arXiv 2607.01418, Jul 2026. https://arxiv.org/abs/2607.01418
  6. Faros AI, Lab vs. Reality: AI productivity study findings, Jul 2025. https://www.faros.ai/blog/lab-vs-reality-ai-productivity-study-findings
  7. Google DORA, State of AI-assisted Software Development 2025, Sep 2025. https://dora.dev/research/2025/dora-report/
  8. Stack Overflow, 2025 Developer Survey (Jul 2025) and survey look-back (Sep 2026). https://survey.stackoverflow.co/2025/ · https://stackoverflow.blog/2026/09/30/getting-ready-for-2026-results-a-look-back-on-developer-survey-findings/
  9. JetBrains, State of Developer Ecosystem 2025, Oct 2025. https://blog.jetbrains.com/research/2025/10/state-of-developer-ecosystem-2025/
  10. GitClear, AI code quality report, Feb 2025. https://gitclear-public.s3.us-west-2.amazonaws.com/GitClear-AI-Copilot-Code-Quality-2025.pdf
  11. Veracode, 2025 GenAI Code Security Report, Jul 2025. https://www.veracode.com/resources/analyst-reports/2025-genai-code-security-report/
  12. Apiiro, 4x Velocity, 10x Vulnerabilities, Sep 2025. https://apiiro.com/blog/4x-velocity-10x-vulnerabilities-ai-coding-assistants-are-shipping-more-risks/
  13. METR, time-horizon benchmark results v1.1 (accessed Oct 2026). https://metr.org/time-horizons/
  14. OpenAI, Why we no longer evaluate SWE-bench Verified, Feb 2026. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
  15. Stanford HAI, 2026 AI Index Report. https://hai.stanford.edu/ai-index/2026-ai-index-report
  16. Gartner, software engineering trends press release, Jul 2025; Magic Quadrant for Enterprise AI Coding Agents, May 2026 (as reported by Virtualization Review). https://www.gartner.com/en/newsroom/press-releases/2025-07-01-gartner-identifies-the-top-strategic-trends-in-software-engineering-for-2025-and-beyond · https://virtualizationreview.com/articles/2026/06/05/ai-firms-push-cloud-giants-from-leaders-quadrant-in-gartner-ai-coding-report.aspx
  17. McKinsey, Unlocking the value of AI in software development. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/unlocking-the-value-of-ai-in-software-development
  18. Google, Cloud Next 2026 keynote remarks on AI-generated code (as reported by The Decoder). https://the-decoder.com/google-says-75-percent-of-its-new-code-is-now-written-by-ai/
  19. Anthropic, Series G announcement, Feb 2026. https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation
  20. Cursor revenue run-rate, Dealroom, Jun 2026. https://dealroom.co/news/134107-cursor-tops-4b-annualized-revenue/
  21. TechCrunch, Y Combinator W25 AI-generated codebases, Mar 2025. https://techcrunch.com/2025/03/06/a-quarter-of-startups-in-ycs-current-cohort-have-codebases-that-are-almost-entirely-ai-generated/
  22. Carta, Solo Founders Report 2025, Dec 2025. https://carta.com/data/solo-founders-report/
  23. US Bureau of Labor Statistics, Occupational Outlook Handbook: Software Developers. https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm
  24. Habituation at the Gate, arXiv 2606.22721, 2026. https://arxiv.org/abs/2606.22721
  25. Kraishan, agent PR security and maintenance study, arXiv 2609.17598, 2026. https://arxiv.org/abs/2609.17598
  26. Purdue agent cost study, arXiv 2609.30725, 2026. https://arxiv.org/abs/2609.30725
  27. GitHub, Octoverse 2025, Oct 2025. https://octoverse.github.com/
  28. GitHub API, star counts for open-source coding-agent repositories, 8 Oct 2026. https://api.github.com/repos/anomalyco/opencode · https://api.github.com/repos/anthropics/claude-code · https://api.github.com/repos/openai/codex · https://api.github.com/repos/google-gemini/gemini-cli · https://api.github.com/repos/OpenHands/OpenHands · https://api.github.com/repos/FoundationAgents/MetaGPT · https://api.github.com/repos/cline/cline · https://api.github.com/repos/block/goose · https://api.github.com/repos/Aider-AI/aider · https://api.github.com/repos/continuedev/continue · https://api.github.com/repos/OpenBMB/ChatDev · https://api.github.com/repos/Kilo-Org/kilocode · https://api.github.com/repos/RooCodeInc/Roo-Code · https://api.github.com/repos/SWE-agent/SWE-agent · https://docs.github.com/en/rest/repos/repos#get-a-repository
  29. McKinsey, Technology Trends Outlook 2026, Sep 2026 (as reported by CIO Dive). https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-top-trends-in-tech
  30. Stripe Engineering, Minions: Stripe's one-shot, end-to-end coding agents, Feb 2026. https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents-part-2
  31. Airbnb Tech Blog, test migration case study, Mar 2025. https://medium.com/airbnb-engineering/accelerating-large-scale-test-migration-with-llms-9565c208023b
  32. AWS DevOps Blog, Amazon Q Developer just reached a $260 million dollar milestone, Aug 2024. https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone/
  33. Uber Engineering, uReview, Aug 2025; The Pragmatic Engineer, How Uber uses AI for development, Mar 2026; Fortune, May 2026. https://www.uber.com/gb/en/blog/ureview/ · https://newsletter.pragmaticengineer.com/p/how-uber-uses-ai-for-development · https://fortune.com/2026/05/22/microsoft-ai-cost-problem-tokens-agents/
  34. Amplifying, State of the Coding Agent Market, Oct 2026. https://amplifying.ai/research/state-of-coding-agents
  35. Takerngsaksiri et al., follow-up fixes after agent merges, arXiv 2609.26847, Sep 2026. https://arxiv.org/abs/2609.26847
  36. SWE-chat, Stanford, arXiv 2604.20779, Apr 2026. https://arxiv.org/abs/2604.20779
  37. Atlassian, State of Developer Experience 2025. https://www.atlassian.com/teams/software-development/state-of-developer-experience-2025
  38. a16z, LLMflation, Nov 2024; Epoch AI, LLM inference price trends. https://a16z.com/llmflation-llm-inference-cost/ · https://epoch.ai/data-insights/llm-inference-price-trends
  39. Anthropic, Claude Code cost documentation. https://code.claude.com/docs/en/costs
  40. Gartner press releases on AI coding costs (Jun 2026) and tiny software teams (Jul 2026). https://www.gartner.com/en/newsroom/press-releases/2026-06-24-gartner-predicts-ai-coding-costs-will-surpass-average-developer-salary-by-2028-as-token-consumption-surges · https://www.gartner.com/en/newsroom/press-releases/2026-07-07-gartner-predicts-60-percent-of-organizations-will-adopt-smaller-software-engineering-teams-by-2029
  41. SaaS Capital, revenue-per-employee benchmarks for private SaaS companies. https://www.saas-capital.com/blog-posts/revenue-per-employee-benchmarks-for-private-saas-companies/
  42. TechCrunch, Wix acquires Base44, Jun 2025. https://techcrunch.com/2025/06/18/6-month-old-solo-owned-vibe-coder-base44-sells-to-wix-for-80m-cash/
  43. Company statements on AI-generated code share, as reported: Google (Apr 2026), Robinhood (Jul 2025), Coinbase (2025), Microsoft (Apr 2025), Uber (May 2026). https://the-decoder.com/google-says-75-percent-of-its-new-code-is-now-written-by-ai/ · https://www.aol.com/robinhood-ceo-says-majority-companys-094801868.html · https://www.theblock.co/news/business/2025-09-04-brian-armstrong-coinbase-ai-369460 · https://techcrunch.com/2025/04/29/microsoft-ceo-says-up-to-30-of-the-companys-code-was-written-by-ai/ · https://fortune.com/company/uber-technologies/earnings/q1-2026/