Specs drive the work
Requirements, acceptance criteria, architecture, and contracts are human-approved and carry permanent IDs. Every change, decision, issue, and regression traces back to the spec it serves.
Run your coding agents inside a governed workflow. Every accepted feature carries its proof, its cost, and its place in the release forecast—and your people keep every consequential decision.
Early-stage evaluation outline. Each capability is demonstrated against agreed acceptance criteria and approved only for that scope. This is product direction—not general availability or a completed roadmap.
F200.ai is the AI Workforce Factory: governed agent teams that adopt company workflows. The Agentic Software Factory is the first specialization, built for the people who commit time and budgets: stable, measured quality, an expected cost and time budget per feature, controlled feature/fix/debt prioritization, and release dates with confidence intervals.
Start with one bounded workflow. Specialized agents turn approved intent into verifiable specifications (spec-driven development, SDD), then build with test-driven development (TDD)—falsifiable acceptance tests written first, implementation until they pass—followed by separate adversarial verification. The same system that runs the work records it, so delivery numbers come from first-hand evidence. People retain product, risk, access, and release responsibility.
Requirements, acceptance criteria, architecture, and contracts are human-approved and carry permanent IDs. Every change, decision, issue, and regression traces back to the spec it serves.
Acceptance tests are written from approved criteria and proven falsifiable—failing before the code exists. A different agent verifies the result adversarially; the implementer cannot edit the tests.
Each feature runs against an expected cost and time budget, with deviations monitored and escalated. A prioritization policy decides what runs next—e.g., features and critical issues first, high/medium issues when capacity or quotas are free, tech debt later, or feature then tech-debt sprints.
Each view is demonstrated against agreed criteria within the selected scope. Every metric shows its source and sample size; forecasts widen when history is thin.
First-pass verification rate, rework, escaped regressions, and gate results from falsifiable tests and adversarial verification—plus feature/fix/debt/regression allocation against your target. Refocus through prioritization lanes or sprint cadences; out-of-policy work is held at dispatch.
Model and runtime usage, retries, and human review and intervention time for each accepted feature, tracked against its expected cost and time budget with deviation alerts.
METR-style time-horizon metrics: the length of task, in human-expert time, agents complete reliably in your codebase—with intervention rate per feature and its trend over time.
Release dates with confidence interval estimates (P95 CI), conditional on budget and concurrency, built from accepted history as it accumulates.
Repository and specification readiness—runnable tests, reproducible environment, acceptance criteria, issue hygiene—with gaps ranked by impact.
The one artifact a reviewer must read: criteria, falsifiable tests, adversarial verification, gates, cost, and effort.
For each required fact, we agree the system of record—the authoritative tool or repository—and who has approved read or write access. Existing tools, terms, roles, approvals, and policies stay in place where valid. Templates fill approved gaps; they do not mandate migration or weaken policy.
A human ownership substrate for specs and requirements, architecture, designs, contracts, and models—everything that needs human review, sign-off, and collaborative work.
Respect backlog types, states, priority, planning, ownership, and acceptance.
Fit branching, testing, review, release, evidence, and responsible operators.
Use approved channels for participation, oversight, and HITL (human-in-the-loop) escalation to an accountable decision-maker.
Map access, required gates, trust boundaries, exceptions, and accountable authority.
The commercial journey and product maturity roadmap are different views. A prepared demo opens the conversation; Discovery defines needs and the isolated greenfield proof of concept (PoC); brownfield Onboarding then prepares the existing product; the Pilot delivers and measures one specified feature through release; Rollout then extends across Agent operations (AgentOps), Delivery cost and usage (FinOps), Governance & Security, and Delivery Intelligence.
A short introduction with a prepared demonstration—not the customer PoC and not proof of the customer environment.
Identify the goal, constraints, decision owners, and an isolated greenfield project for the customer-specific PoC.
Evaluate the workflow on an isolated greenfield project—not the customer’s existing repository and not brownfield onboarding.
Reconstruct the existing repository, documents, backlog, tools, delivery process, collaboration, and policies. Start from a Readiness Report, address feature-relevant gaps, record decisions and architecture decision records, and specify one feature, acceptance criteria, and release.
Start only after completed Onboarding is approved: acceptance criteria are approved, feature-relevant gaps are addressed, required decisions are recorded, and the release path is authorized. Implement that one feature in the existing product with TDD, verify it independently, follow customer approvals, and produce the release, the acceptance package, and a measured delivery result with its cost observation.
Review the measured result, agree commercial terms, then roll out across Agent operations (AgentOps), Delivery cost and usage (FinOps), Governance & Security, and Delivery Intelligence—or stop.
The acceptance package is the concise review record for the customer’s decision: the specified change, the evidence and result, security and compliance posture, risks and trade-offs, measured spend, human effort and autonomy, ownership, and the decision requested.
Acceptance, constraints, authority, exclusions, and material decisions.
Stage by stage: greenfield PoC evidence; Onboarding baseline, gaps, and decisions; Pilot tests, adversarial verification, approvals, and the exact authorized release. Deployment is a separate approved event with health and rollback evidence.
Security findings, access and trust boundaries exercised, data flows and model/provider use, policy gates passed, and exceptions with their approvers.
Risk register with likelihood, impact, owner, and mitigation; trade-offs taken; open gaps and unknowns; impact on scope, date, and budget; release status; and the decision requested.
Model and runtime spend, human review and intervention effort, and the autonomy level reached, against budget—marked known, estimated, or unknown, never false precision.
Who decides and who operates; runbooks, monitoring, and support handoff; and what requires a separately agreed expansion.
This is a planned product sequence, not completed status or promised dates—a separate view from the customer journey above. Each tier builds on the one before, from augmenting individual engineers to a security-hardened, governed factory.
Engineer’s augmentation with factory control-plane & local runners
AI workforce with collaborative integrations (Jira, Linear, Slack, GitHub)
Automated conveyor for S-SDLC workflows with agent fleet orchestration
Security-hardened factory with governance & compliance controls
Commercial and technical details are settled for the specific evaluation—not implied by a generic package.
No forced migration is part of the evaluation. We map your five domains, preserve valid sources and practices, and propose changes only where stakeholders agree a gap exists.
Your authorized people own product intent, acceptance, the system of record and write authority for each required fact, access, policy exceptions, release, and consequences. Agents work within delegated bounds; silence is not approval.
Large language model (LLM) work uses approved, subscription-backed customer clients. Provider terms, model choice, tool access, data handling, and any third-party processing require case-specific review and approval. We do not claim that all code or data remains on one machine.
Evaluation fees, included support, and budget limits are agreed before work. Customer AI subscriptions and third-party or infrastructure costs remain separate unless the written scope says otherwise. No fixed price or guaranteed savings is implied here.
Forecasts are built from accepted features as they accumulate, conditional on budget and concurrency. Each forecast shows its sample size and calibration; with little history the interval is wide. A Pilot delivers one feature, so it starts the history rather than proving a forecast.
No availability claim is made by this outline. It describes an early-stage product direction and a proposed evaluation structure. Each capability, assurance, and integration must be demonstrated and approved within the selected scope.
Bring one workflow, the person who owns its outcome, and one result worth measuring. Discovery follows only if the prepared Demo shows a useful fit.
F200.ai · AI Workforce Factory · Software-first product direction. This overview is conversation material, not an offer, contract, certification, or availability commitment.