Agentic Software Factory · for delivery teams

Predictable delivery from an AI workforce. Start with one feature.

Run your coding agents inside a governed workflow. Every accepted feature carries its proof, its cost, and its place in the release forecast—and your people keep every consequential decision.

Early-stage evaluation outline. Each capability is demonstrated against agreed acceptance criteria and approved only for that scope. This is product direction—not general availability or a completed roadmap.

The operating idea

The accepted feature is the unit of governance.

F200.ai is the AI Workforce Factory: governed agent teams that adopt company workflows. The Agentic Software Factory is the first specialization, built for the people who commit time and budgets: stable, measured quality, an expected cost and time budget per feature, controlled feature/fix/debt prioritization, and release dates with confidence intervals.

Today’s friction

  • AI output is fast, but whether it meets the requirement is unproven.
  • Spend arrives as seats and tokens, not cost per feature.
  • Features, fixes, and technical debt compete for the same agents, unseen—and refocusing means re-sorting every backlog by hand.
  • Release dates rest on history that agent adoption just changed.

Intended outcome

  • Acceptance tests are falsifiable and adversarially verified against acceptance criteria.
  • Each accepted feature carries its cost—expected cost & time budget, with deviation monitoring.
  • Feature/fix/debt prioritization is set by policy and enforced at dispatch; refocusing is a policy change, not a backlog rewrite.
  • Release forecasts account for uncertainty with confidence interval estimates (P95 CI).
How it works

Intent, Quality, Control.

Start with one bounded workflow. Specialized agents turn approved intent into verifiable specifications (spec-driven development, SDD), then build with test-driven development (TDD)—falsifiable acceptance tests written first, implementation until they pass—followed by separate adversarial verification. The same system that runs the work records it, so delivery numbers come from first-hand evidence. People retain product, risk, access, and release responsibility.

01 · INTENT (SDD)

Specs drive the work

Requirements, acceptance criteria, architecture, and contracts are human-approved and carry permanent IDs. Every change, decision, issue, and regression traces back to the spec it serves.

02 · QUALITY (TDD)

Tests prove the spec

Acceptance tests are written from approved criteria and proven falsifiable—failing before the code exists. A different agent verifies the result adversarially; the implementer cannot edit the tests.

03 · CONTROL

On time and on budget

Each feature runs against an expected cost and time budget, with deviations monitored and escalated. A prioritization policy decides what runs next—e.g., features and critical issues first, high/medium issues when capacity or quotas are free, tech debt later, or feature then tech-debt sprints.

What delivery leaders get

Numbers you can steer by.

Each view is demonstrated against agreed criteria within the selected scope. Every metric shows its source and sample size; forecasts widen when history is thin.

Stable delivery, measured quality

First-pass verification rate, rework, escaped regressions, and gate results from falsifiable tests and adversarial verification—plus feature/fix/debt/regression allocation against your target. Refocus through prioritization lanes or sprint cadences; out-of-policy work is held at dispatch.

Cost & effort per accepted feature

Model and runtime usage, retries, and human review and intervention time for each accepted feature, tracked against its expected cost and time budget with deviation alerts.

Agentic autonomy level

METR-style time-horizon metrics: the length of task, in human-expert time, agents complete reliably in your codebase—with intervention rate per feature and its trend over time.

Release forecast

Release dates with confidence interval estimates (P95 CI), conditional on budget and concurrency, built from accepted history as it accumulates.

Readiness Report

Repository and specification readiness—runnable tests, reproducible environment, acceptance criteria, issue hygiene—with gaps ranked by impact.

Acceptance package

The one artifact a reviewer must read: criteria, falsifiable tests, adversarial verification, gates, cost, and effort.

Meet you where you are

Adoption across five connected domains.

For each required fact, we agree the system of record—the authoritative tool or repository—and who has approved read or write access. Existing tools, terms, roles, approvals, and policies stay in place where valid. Templates fill approved gaps; they do not mandate migration or weaken policy.

01

Documentation Management

A human ownership substrate for specs and requirements, architecture, designs, contracts, and models—everything that needs human review, sign-off, and collaborative work.

02

Project Management

Respect backlog types, states, priority, planning, ownership, and acceptance.

03

Software development lifecycle

Fit branching, testing, review, release, evidence, and responsible operators.

04

Human–Agent Collaboration

Use approved channels for participation, oversight, and HITL (human-in-the-loop) escalation to an accountable decision-maker.

05

Governance & Security

Map access, required gates, trust boundaries, exceptions, and accountable authority.

Customer journey

From prepared demo to measured release.

The commercial journey and product maturity roadmap are different views. A prepared demo opens the conversation; Discovery defines needs and the isolated greenfield proof of concept (PoC); brownfield Onboarding then prepares the existing product; the Pilot delivers and measures one specified feature through release; Rollout then extends across Agent operations (AgentOps), Delivery cost and usage (FinOps), Governance & Security, and Delivery Intelligence.

01 · INTRO + DEMO

See the prepared flow

A short introduction with a prepared demonstration—not the customer PoC and not proof of the customer environment.

02 · DISCOVERY

Define needs and PoC scope

Identify the goal, constraints, decision owners, and an isolated greenfield project for the customer-specific PoC.

03 · GREENFIELD PoC

Evaluate the Factory workflow separately

Evaluate the workflow on an isolated greenfield project—not the customer’s existing repository and not brownfield onboarding.

04 · ONBOARDING

Prepare the brownfield product

Reconstruct the existing repository, documents, backlog, tools, delivery process, collaboration, and policies. Start from a Readiness Report, address feature-relevant gaps, record decisions and architecture decision records, and specify one feature, acceptance criteria, and release.

05 · BROWNFIELD PILOT

Deliver and measure one specified feature

Start only after completed Onboarding is approved: acceptance criteria are approved, feature-relevant gaps are addressed, required decisions are recorded, and the release path is authorized. Implement that one feature in the existing product with TDD, verify it independently, follow customer approvals, and produce the release, the acceptance package, and a measured delivery result with its cost observation.

06 · ROLLOUT

Expand across operations

Review the measured result, agree commercial terms, then roll out across Agent operations (AgentOps), Delivery cost and usage (FinOps), Governance & Security, and Delivery Intelligence—or stop.

PoC and Onboarding are separate. The PoC is an isolated greenfield evaluation. Onboarding is brownfield preparation: observations from existing code are not automatically intent, feature-relevant gaps and decisions must be resolved, and no migration or policy change is forced. Completed Onboarding gates the Pilot.
Decision-ready output

A concise acceptance package.

The acceptance package is the concise review record for the customer’s decision: the specified change, the evidence and result, security and compliance posture, risks and trade-offs, measured spend, human effort and autonomy, ownership, and the decision requested.

Agreed intent and scope

Acceptance, constraints, authority, exclusions, and material decisions.

Evidence and result

Stage by stage: greenfield PoC evidence; Onboarding baseline, gaps, and decisions; Pilot tests, adversarial verification, approvals, and the exact authorized release. Deployment is a separate approved event with health and rollback evidence.

Security & Compliance posture

Security findings, access and trust boundaries exercised, data flows and model/provider use, policy gates passed, and exceptions with their approvers.

Risk and decision brief

Risk register with likelihood, impact, owner, and mitigation; trade-offs taken; open gaps and unknowns; impact on scope, date, and budget; release status; and the decision requested.

Spend and effort observations

Model and runtime spend, human review and intervention effort, and the autonomy level reached, against budget—marked known, estimated, or unknown, never false precision.

Ownership and Operations

Who decides and who operates; runbooks, monitoring, and support handoff; and what requires a separately agreed expansion.

Planned product roadmap

From engineer’s toolkit to enterprise factory.

This is a planned product sequence, not completed status or promised dates—a separate view from the customer journey above. Each tier builds on the one before, from augmenting individual engineers to a security-hardened, governed factory.

1

OSS Toolkit

Engineer’s augmentation with factory control-plane & local runners

2

Cloud Agents

AI workforce with collaborative integrations (Jira, Linear, Slack, GitHub)

3

Factory

Automated conveyor for S-SDLC workflows with agent fleet orchestration

4

Enterprise Factory

Security-hardened factory with governance & compliance controls

Trust & adoption FAQ

Clear boundaries before access.

Commercial and technical details are settled for the specific evaluation—not implied by a generic package.

Does this replace our tools or process?

No forced migration is part of the evaluation. We map your five domains, preserve valid sources and practices, and propose changes only where stakeholders agree a gap exists.

Who remains responsible?

Your authorized people own product intent, acceptance, the system of record and write authority for each required fact, access, policy exceptions, release, and consequences. Agents work within delegated bounds; silence is not approval.

Where do models run and where does data flow?

Large language model (LLM) work uses approved, subscription-backed customer clients. Provider terms, model choice, tool access, data handling, and any third-party processing require case-specific review and approval. We do not claim that all code or data remains on one machine.

How is spend handled?

Evaluation fees, included support, and budget limits are agreed before work. Customer AI subscriptions and third-party or infrastructure costs remain separate unless the written scope says otherwise. No fixed price or guaranteed savings is implied here.

Can you forecast our release date?

Forecasts are built from accepted features as they accumulate, conditional on budget and concurrency. Each forecast shows its sample size and calibration; with little history the interval is wide. A Pilot delivers one feature, so it starts the history rather than proving a forecast.

Is this generally available?

No availability claim is made by this outline. It describes an early-stage product direction and a proposed evaluation structure. Each capability, assurance, and integration must be demonstrated and approved within the selected scope.

A practical next conversation

Start with an intro + prepared Demo.

Bring one workflow, the person who owns its outcome, and one result worth measuring. Discovery follows only if the prepared Demo shows a useful fit.

F200.ai · AI Workforce Factory · Software-first product direction. This overview is conversation material, not an offer, contract, certification, or availability commitment.