Overhead harbour chart titled Shiftyard, with a dock building, four terminal-equipped berths, and a magenta route from work cards through a gatehouse to a harbour-mouth junction.Work follows a shared course through the Shiftyard harbour.

Coding agents are good enough now that the hard part is no longer getting one to write a change. The hard part is running ten of them at once, on your own subscriptions, against real repositories, without one of them force-pushing to main, leaking a credential into a prompt, or quietly merging a pull request nobody reviewed.

Shiftyard is the software we built to do that for ourselves. It is a daemon that runs unmodified coding-agent CLIs (Claude Code and Codex today) on machines you own, and wraps every one of them in a queue, leases, a policy gate, budgets, verification and a merge train. Since 5 October 2026 our own engineering work, including most of Shiftyard's own commits, has run on it.

This post is a long, technical tour: the architecture, how a command gets allowed or refused, how two agents are kept from writing over each other, how a change travels from a queue item to a merged pull request, how the memory subsystem and its models are being selected, what we test and what we have measured. It also says plainly what is built, what is still in review, and what is only planned.

Status. Shiftyard is pre-release: there is nothing to install yet. The code is licensed Apache-2.0 and the public repository is coming soon. Everything below describes the code as of early October 2026 on macOS hosts.

1. Why we built it

A single coding agent in a terminal is a solved workflow. A team of them is not. Once you run more than a handful in parallel you meet the same failures over and over:

  • Shared state. Two agents claim the same ticket, or both rebase the same branch, or one writes after its claim has expired.
  • Unbounded actions. An agent that can run git push --force, terraform destroy or gh pr merge will, eventually, run it at the wrong moment.
  • Credentials. Vendor logins, cloud keys and tokens drift into prompts, logs and model context.
  • Verification theatre. An agent says the tests pass. Whether they ran, on which commit and with which files changed is another question.
  • Capacity. Subscriptions have usage windows, laptops have RAM and thermal limits, and a resumed session costs less than a fresh one.
  • Memory. Every new session re-derives what the last one already worked out.

We had been running an earlier internal orchestration platform for this. Shiftyard is its rewrite: one Rust daemon per host, with every backend service native inside it, built around the lessons from operating the first one.

2. Principles

Four principles are written into the project's README, and most design decisions trace back to them:

  1. Your credentials stay on your machine. You sign in to each agent through its vendor's own login flow. Shiftyard never reads, stores or forwards those credentials.
  2. One agent by default. A builder plus QA pair is an opt-in verification step, not the baseline.
  3. Deterministic enforcement. Policy gates, sandboxing and budgets run on the host. They do not depend on the agent's cooperation.
  4. Works offline. Local work never depends on the cloud.

Two engineering rules follow from the third principle and shape the whole codebase:

  • Pure domain cores. The gate, leases, the work state machine, flows, budgets, capacity, the balancer and the model router are pure crates: no async runtime, no I/O and no clock. Time is passed in. That makes every rule property-testable and replayable.
  • One writer. A single SQLite writer thread owns every SQL statement; async code never writes to the database directly. Concurrency is three sanctioned patterns only: actor tasks with channels, the writer thread, and cancellation tokens for shutdown.

3. One host, end to end

Host enclosure containing a daemon with API, scheduler, gate, store, land, and egress modules above three sandboxed terminal seats with lock icons and arrows to the gate, plus one egress exit.One host contains the daemon, sandboxed seats, and a single egress path.

A yard is one install: the daemon on one host and everything it runs. A crew is a named group of pods and seats with one purpose. A seat is a durable position with a role and an address; the agent session in it can change.

A macOS host contains a single daemon with eight modules and isolated seats. Operator traffic enters the API over loopback, hooks wait on the policy gate, bare-mirror objects feed each clone, and only redacted text exits through the egress sanitizer.One host contains the daemon, isolated seats, and a single sanitized path out.

The daemon exposes two listeners and nothing else: an operator listener on loopback (admin token or session cookie, browser defences, the embedded web UI, server-sent event streams) and a Unix socket for seats (per-seat tokens, daemon user only). There is deliberately no non-loopback listener, no TLS termination and no CORS. The local API is a typed RPC service of 94 methods, served as JSON or binary, with every schema change gated by a breaking-change check.

4. Seats: running unmodified agent CLIs

Shiftyard does not fork or patch the agents. It runs the vendor's own CLI and controls it from the outside, through the same surfaces a careful human would use: the terminal, the CLI's hook system and its configuration.

The launch plan. Every launch, resume and wake recomputes a launch plan as a pure function of the crew spec, the policy and the seat's state. It is never stored as truth; it is written to disk at mode 0600 just before launch, for the record.

The PTY holder. The agent runs inside a small holder process that owns its pseudo-terminal. The daemon talks to the holder over an authenticated socket with a length-prefixed protocol that supports the current and previous version, so the daemon can be upgraded while seats keep running. Input is typed only as named keys or bracketed paste.

Reading the screen. Detection reads the rendered screen, not the byte stream. We chose the terminal model with a spike that replayed real recorded sessions against a reference emulator (section 14): coding-agent TUIs draw spaces with cursor moves and put non-breaking spaces after their prompt glyph, so pattern-matching raw output is not reliable.

Fenced stops. Every stop is fenced to a launch id and a generation. Settlement proves that the agent's process id, start time and process group are gone before a seat is considered stopped, so a late signal can never hit a newer launch.

Claude Code. Each seat gets its own settings file with hooks and a status line, its own account directory and a private MCP configuration loaded in strict mode, so account-level MCP servers never leak into a seat. Hooks call back into the daemon over a private socket, bound to the launch nonce. PreToolUse waits for the gate's answer and fails closed. Until a session-start event proves which session is running, every tool call is denied. A usage-limit menu is only ever answered "stop and wait". A resume never silently becomes a fresh start; only the recycler starts fresh, deliberately, with a recovery packet.

Codex. Each seat gets its own CODEX_HOME and a per-seat profile with hooks that route session start, prompt submit and tool calls to the gate. A deny or ask returns exit code 2. The account's own configuration is checked against an allowlist of keys pinned to one CLI version; account-level hooks, rules or skills refuse the launch rather than being silently stripped. Dangerous bypass flags are banned outright, and a resume first checks that the rollout exists and that no other writer holds it.

Credentials. The account registry only probes identity: which email a Claude account belongs to, and whether a keychain item exists. The Codex driver is not allowed to open the credentials file at all.

Other agents. An OpenCode driver exists (two sandboxed processes: a pinned server and an attached TUI, every permission request routed to the gate, and the daemon re-reading the agent's own store every second to detect rules it did not set). It is in review, not yet composed into the running daemon. Support for further CLIs exists only as offline interface probes.

5. The policy gate

Terminal commands pass through parse and classify stages into a rule ladder that branches to a large navy DENY bin, a medium slate ASK bin, and a small magenta ALLOW bin.The policy gate lets the most restrictive rule determine the outcome.

Every shell command an agent wants to run goes through the gate before it runs. The gate is a pure function: no clock, no store, no network and no model.

Five policy stages parse command text, classify typed action kinds, apply eight first-match rules, combine outcomes as DENY over ASK over ALLOW, and atomically record authority and consume a grant in the hash-chained audit.The policy gate fails closed and records decisions with their authority.

A few details matter more than they look:

  • Obfuscation fails closed. Command substitution the parser cannot resolve, continuation tricks and a handful of grammar edge cases are marked opaque on purpose. Only a person can allow an opaque action.
  • Protected refs. main and master are always protected, in any spelling and case. Writes into .git/ or hook directories need a person's approval of the exact command.
  • A model may only tighten. If a model is consulted at all, it can turn an Allow into an Ask. Nothing ever becomes an Allow because a model said so.
  • Pinned policy. Policies are YAML with tightening overlays, canonicalised to JSON and identified by their SHA-256. Each launch is pinned to one policy hash. An invalid document fails closed; there is never a fallback policy.
  • Panics are denials. A classifier panic is caught and becomes Deny.
  • The daemon being down is a denial. If the hook cannot reach the daemon, every tool call is refused, reads included. Approvals are never auto-answered.

The code states the gate's limits in its own documentation, and so do we: the gate is defence in depth against mistakes and runaway agents, not a security boundary. The controls that hold against a hostile agent are server-side (scoped credentials, branch protection) and the OS sandbox. The gate exists so that the common failure, an honest agent doing the wrong thing at the wrong time, is stopped deterministically and recorded.

6. Leases and fences

Two seat lanes cross a timeline fence marked 7 to 8; Seat A's token 7 write bounces back, while Seat B's token 8 write passes through.Fence tokens reject stale writes and admit the current lease holder.

Anything two seats could fight over is a leased resource: a work item, a slot in a heavy-job pool, a branch, a step in a flow, the merge train of a repository, the single writer of an issue-tracker project, or a claimed path in a repository.

A lease alone is unsafe. A holder can stall past its expiry, wake up and write anyway. So every resource also carries a fence: a number that only grows. Each grant and each revoke moves it on, and every guarded write checks the writer's fence inside the same transaction as the write.

Seat A receives fence 7 and stalls; its lease expires, Seat B receives fence 8, and the store refuses A’s stale write while accepting B’s current write. External writers use HMAC tokens and retain each resource’s highest fence.A newer fence prevents an expired holder from writing stale work.

The rules are one pure function, decide(view, command, now), over six commands: grant, renew, release, validate, expire and revoke. The store applies the resulting changes inside the caller's write transaction. A few constants keep it honest: no lease lives longer than 24 hours, a holder may have at most 8 live leases, and a grant is refused if the clock reads more than 5 seconds earlier than the newest grant, so a clock that jumps backwards can neither mint grants nor shorten a lease.

Holders are typed (seat, job, flow, host), and cleanup for one type never touches another type's leases. That rule exists because the earlier platform once let a seat's leak guard free a background job's lease.

For writers outside the database transaction, the daemon issues a signed, expiring fence token keyed by a host secret that never appears in an agent's environment or logs. The receiver checks the signature in constant time first, then refuses any fence lower than the highest it has accepted for that resource.

Every lease implementation, the local engine today and any future broker, must pass the same shared conformance fixtures: stale renew and release, revoking a stale write, idempotent regrant, double expiry sweeps, conflicting grants and validation after expiry.

7. The work queue

Work items move through an explicit state matrix. There are four live states and five terminal ones, and terminal states refuse every command.

From Command To
Pending Claim Claimed
Pending Block / Handoff / Deny / Cancel as named
Claimed Start Running
Claimed Complete / Fail Done / Failed
Claimed Block / Handoff / Deny / Cancel as named
Running Complete / Fail Done / Failed
Running Block / Handoff / Cancel as named
Blocked Unblock Pending, or Claimed if the prior claimant can take it back
Blocked Claim / Block / Handoff / Deny / Cancel as named

An exhaustive test walks every state against every command. Claims are leases with fences, so a seat transitions its own item by presenting its fence. When an item closes, items blocked directly on it are released in the same transaction.

Waiting without polling. Agents are bad at waiting. Shiftyard bans polling outright: a seat that needs CI to finish blocks its item on a canonical external reference (a commit's checks, a pull request, a CI run) with a wake-after between 5 minutes and 24 hours, and ends its turn. The daemon wakes it when the reference changes, with a slow fallback wake if nothing arrives. A wake is not a result: the seat re-reads the evidence when it wakes.

Idle dispatch. An idle seat is offered its next pending item addressed to it, most urgent first and then oldest, once per idle window. The seat claims it itself. Finishing work comes before starting work.

Work-in-progress limits. New build claims wait while a repository already has 15 open pull requests from the factory.

8. Verification and the land train

Pull request cards on train wheels follow a magenta track through verify, review, approve, and merge signals to a main station, with queued cars stopped at the red verify signal.The land train carries pull requests through checks toward main.

Two mechanisms stand between an agent saying "done" and a change reaching main.

Stop-hook verification. A seat may finish its turn only once the repository's declared checks pass. The check declaration is read from the daemon's own mirror at a base commit the daemon resolved, never from the seat's working copy. The checks run in a fresh checkout built from the work's tree, with protected paths (tests, fixtures, CI configuration) taken from the base, and with a cleared environment. So an agent cannot make its turn pass by editing the tests or the check declaration. A pass lets the turn end; a failure blocks it and hands back the redacted output tail; after repeated refusals a person decides. Verdicts currently record isolation = weak, because the separate verify user does not exist yet.

The land train.

A two-row pipeline follows a queue item through a fenced claim, seat edits, stop-hook verification, PR creation, risk-tier reviews, SHA-bound checks, landability, person-approved LAND NOW, a pinned-head merge driver, and main; failed verification blocks the turn and a failed release verdict freezes the train.Every merge passes verification, reviews, current-head checks, and a person-approved land batch.

Details that came from real incidents:

  • Verdicts are bound to the exact head commit. A verdict on an older head does not count.
  • Rebases. A verdict is carried across a rebase only when the base, the patch-id and the set of changed files are all the same. Otherwise it is superseded and the PR gets one re-review per kind. CI results never carry.
  • Risk tiers. Tiers come from path rules. An unclassified PR is normal; unknown files are read as infrastructure. An infrastructure batch is always size one.
  • Only the head the verify gate admitted can land. The merge driver only touches pull requests the engine opened, pins the head SHA on the merge and never uses an admin bypass. It runs every ten minutes and shortly after any verdict, approval or green check.

9. Scheduling, capacity, routing and budgets

Liveness from evidence. The wake core never stores "alive" as a flag. Liveness is computed from evidence: the holder, the process, the screen, the hooks. Every launch is a resume; there is no fresh launch except through the recycler. A permanent-failure breaker stops a seat that keeps failing to launch.

Housekeeping that is on by default.

Behaviour What it does
limit holds a seat that hits a usage limit has its wakes held until the reset; two agreeing seats open an account-wide limit episode
stuck alarms a busy seat with no progress raises an item to its coordinator
orphan adoption after a daemon restart, adopt a live agent whose holder survived, after proving identity (fenced token, kernel peer uid, holder executable, pid and start time, plan hash)
reaper sleep idle seats (10 minutes for unit seats, 30 for core seats) after saving their work-in-progress to a ref and snapshotting the screen
recycler start a fresh session for an implementer at 60% context or after a finished item, seeded with a recovery packet; others are compacted
idle dispatch offer an idle seat its next item

Capacity. The working seat target is floor(min(ram_seat_cap, cpu_seat_budget) × thermal_factor). The daemon samples the host every 15 seconds, keeps a 2 GiB memory reserve plus 1 GiB per new seat and a disk reserve, and gates queued resumes on CPU below 85%. Heavy jobs (builds, test suites) take a slot from a leased pool through a governor command that runs the job in its own process group, renews the slot at a third of its remaining time and releases it when the job ends. A heat guard that renices builds, caps concurrent build trees and backs off on thermal pressure is designed and in review.

Routing tiers and floors. Models are grouped into three tiers. A crew may lower a tier; only the policy may raise one. Each stage of work has a floor:

Stage Minimum tier Effort
implement / fix (small, medium) 2 medium
implement / fix (large or unknown) 3 medium
QA and code review 2 medium
security review 3 high
security delta review 2 high
coordination 3 medium
design 3 high

Tier 3 never runs at low effort.

Budgets and usage. Budget admission is pure fixed-point arithmetic. Token use and list price are read from each CLI's own local records into a usage ledger, and metrics can arrive over a loopback OpenTelemetry receiver.

The balancer. A pure engine decides where seats should run across accounts: hard filters that fail closed, then ranking by headroom before the next reset, the cheapest suitable role chain and slack; failover at a cold boundary and fail-back after reset; and caps learned for windows a vendor does not publish. Today the balancer records its moves; applying them is in review.

Model gateway. For components that call models directly, the daemon serves a loopback gateway that speaks the OpenAI Responses, OpenAI chat-completions and Anthropic messages protocols, walks a fallback chain with per-entry cooldowns, classifies errors, meters usage into the ledger and redacts on the way out.

10. Isolation and egress

Sandbox. On macOS every seat runs under a per-seat, deny-by-default Seatbelt profile (derived from the Apache-2.0 Codex sandbox profiles). It hides other seats' state, allows only the seat's own terminal, restricts loopback by address family and denies terminal input injection. Applying the profile fails closed, and an escape test suite runs against it. Two limits are documented rather than hidden: Seatbelt does not stop one process from reading another's arguments through the kernel, and there is no Linux sandbox yet.

Agent OS identity. A small root-installed helper can run agents under their own OS user, with unsafe code confined to one module. The running daemon still launches seats as the same user today.

Workspaces. Each seat has its own clone of each repository, borrowing objects from a daemon-owned bare mirror. Work in progress is saved to a per-seat ref, retired clones are parked and pruned after 30 days, and build directories are warmed with copy-on-write.

Egress. There is exactly one way text leaves a host. A redacted-text type can only be constructed by the sanitizer, and every transport accepts nothing else. Deny terms are matched in a folded form, the egress log keeps only hashes and counts, never content, and a missing or empty deny list means nothing leaves at all. An operator can preview exactly what would leave before it does.

11. Storage and audit

  • One writer thread plus a pool of read-only connections.
  • Two databases. The engine database, and an attached, append-only, hash-chained audit database. An audit command recomputes the chain end to end.
  • Migrations are forward-only and checksummed, with a backup before every run.
  • The outbox. Side effects (GitHub calls, issue updates, notifications) are written as outbox rows in the same transaction as the state change and executed by an effect runner, so a crash cannot produce a merged PR without a record or a record without its effect.
  • Redaction before storage. Credential shapes are replaced before any text is stored.
  • Secrets. Daemon-held secrets are sealed with XChaCha20-Poly1305 bound to the secret's name, with the data key in the OS keyring. The secret type cannot be formatted and is wiped on drop.
  • Authority records can only be written through a sealed key type, and a dependency policy check bans every other crate from depending on it.

12. Memory: index, embeddings and governance

Dashed source icons for git, PRs, issues, and sessions connect to solid chunker and embed components, vector dots, a chained ledger book, and a dashed search magnifier.The memory pipeline connects chunking, embeddings, and a hash-chained ledger with planned sources and search.

Agents forget. Our team has been running a separate memory service for months: it indexes our repositories, git history, every agent session and a ledger of remembered decisions, and serves all of it over MCP. At its last backup it held 179,037 indexed chunks. It is written in Go and uses a local embedding model through a separate model server.

On 6 October we decided that memory should not be a separate product. It becomes native code inside the Shiftyard daemon. The building blocks are in the repository; composing them into the daemon and cutting over is the work of the coming days. This section separates the two.

Solid built memory blocks include chunking, embedding, the index store, and a governed append-only ledger; dashed planned blocks include sources, hybrid search, and MCP tools. A model-selection gate lists throughput, cosine, ranking, and licence criteria.Memory processing and governance are built while ingestion and retrieval remain planned.

Chunking. We ported the per-language regex rules of the existing service rather than switching to tree-sitter. Pure Rust, the same languages and output that can be diffed against the reference. Code is split at top-level definitions together with their leading comments and attributes; prose is split at headings outside code fences; small spans are merged and large ones split on line or paragraph boundaries; cuts move back to a UTF-8 boundary; every piece of an oversized code span is capped at 4,000 bytes. A parity test checks spans, symbols and imports against the reference implementation's output; on a 32-file test corpus the outputs match except for two documented deviations. A tree-sitter chunker is a later decision, gated on beating the regex baseline on code-scope retrieval.

Embeddings. The model layer defines Embedder, Reranker and Classifier interfaces and one shared post-processing pipeline: an 800-character input cap, the model card's query and document prefixes, Matryoshka truncation, L2 normalisation and an int8 form with sign bits. A model id names a vector space (name, precision, dimensions and a hash prefix), and vectors from different ids are never mixed.

Models are open-weight GGUF files only, run in process through llama.cpp: Metal on macOS, CPU elsewhere, GPU by default with an automatic CPU fallback whose reason is recorded. One thread owns a loaded model. User text is tokenised with special-token parsing switched off, so a document cannot inject control tokens. We chose not to use ONNX Runtime.

Model supply chain. A model manifest pins a 40-character commit and a SHA-256, accepts only Apache-2.0 or MIT licences with no override, and accepts GGUF only, never pickle. Provisioning asks for consent with the size and licence shown, downloads with a size cap and resume, verifies the file before every load, works offline once cached and uses file locks for fetch and use. One candidate embedding model was dropped from consideration purely on its licence terms.

Governed memory. Remembered facts and decisions live in an append-only, hash-chained lifecycle ledger: approve, hold, supersede, retract, quarantine and restore, each atomic with its ledger event. Imports are idempotent and keep provenance. A write to a shared scope stays pending until an owner approves it.

Cutover plan. Backup (done 6 October), a four-model quality and speed evaluation, a parity proof, composition into the daemon, a 48-hour shadow run, then the switch, one seat first and then all, and finally the old service read-only and retired. The target window is 8 to 15 October. The rule written into the plan: if no model meets both the quality and the speed bar, we report the measured results; we never lower either bar.

13. Models and training

Shiftyard trains no model weights. That is an explicit non-goal, not an omission. Training models, or using another model's outputs as labels, is out of scope for the memory work. Hosted judgment results are never used as training data. What we do instead is select open models against fixed gates, and optimise text (guidance and skills) against held-out evaluations.

Selecting the embedding model. The candidates are open-weight embedders in the 0.1 to 0.6 B parameter range: Qwen3-Embedding-0.6B, nomic-embed-text-v1.5, the BGE family (bge-m3, bge-small-en-v1.5) and IBM Granite embeddings. A candidate passes only if all of these hold:

Gate Bar
throughput median warm chunks/s at 800 characters above the current baseline, and above 15
numerical agreement mean cosine against the same model under the reference runtime ≥ 0.99, minimum ≥ 0.97
retrieval quality file-level MRR@10 ≥ baseline − 0.01
licence Apache-2.0 or MIT

Cutover itself has a stricter, statistical gate: a paired bootstrap with 10,000 resamples on at least 150 queries, whose 95% confidence interval lower bound must be ≥ −0.02 on both MRR@10 and weighted recall@10, with search p95 at or under 293 ms and a daemon footprint at or under 78 MB excluding the model, at 15 seats. Quantisation has its own gate (cosine ≥ 0.99, ΔMRR ≤ 0.01, classifier ΔF1 ≤ 1 point). A reranker is optional and must earn its place on nDCG@10.

What we have measured so far.

Measurement Result Status
reference service, M2, 3,000-char inputs 6.1 chunks/s measured
reference service, M2, 1,500-char inputs 14.3 chunks/s measured
reference service, M2, 800-char inputs 15–28 chunks/s measured
reference service, pooled file-level MRR@10 at 800 chars 0.529 measured
native, Qwen3-Embedding-0.6B Q8_0 (1024-d), batch 32, M2 + Metal 11.3 chunks/s preliminary: laptop under load, not a clean run
reranking in the reference service hit@1 −0.085 measured: it hurt, so it is off by default

The native number is below the baseline range, and we are publishing it anyway: it was taken on a loaded laptop and is not a gate result. Performance claims in the codebase must be measured in a fresh process, on fixed inputs, after a warm-up, as the median of several runs. The negative reranking result is the reason the reranker is opt-in and gated.

Experimental: automatic guidance and skill learning. On an unmerged branch we are testing a loop that improves the text agents run on (their guidance files and skills) without training anything:

  • A local Qwen2.5-1.5B-Instruct model (Q4_K_M GGUF, about 1.1 GB) runs in a daemon-managed llama.cpp child process. API spend is zero, each run has a 10-minute budget and makes at most three changes.
  • Runs happen daily (guidance in the morning, skills an hour later) and after outcome events.
  • A candidate must strictly improve on the training set, then pass one sealed, paired, held-out evaluation, then a person reviews it before it is activated.
  • Retrieval inside the loop uses BGE-small-en-v1.5 (Q8, 384-d) for a shortlist of 8, a Jina turbo reranker, and BM25 plus dense retrieval fused with reciprocal rank fusion (k = 60).

The measurements so far are deliberately small and do not establish production quality. On 10 English and 10 Ukrainian queries, top-1 retrieval was 80% and 90% with query translation, 20% for Ukrainian without translation, and 70% for English before fusion. On 20 fresh English queries it moved from 65% to 70%. In screening, four smaller models (Qwen2.5-0.5B, Qwen3-0.6B, Qwen3.5-0.8B and SmolLM2-360M) each failed at least one of five guidance checks. No change has yet earned promotion.

Advisory judgments. Four typed judgments help the operator: queue priority, duplicate ticket, a seat stuck in a loop, and pull-request scope drift. Each has a firm probability band (for example p ≥ 0.9 for a duplicate). They are advisory only. A judgment may add a note, a badge or a display order; it may never grant, block, merge or reorder a queue, and "no answer" must behave exactly like "off". The backend is hosted or off today, and a local decider that shadows the hosted one is planned.

14. Testing and measurements

The codebase on main at the time of writing:

Metric Value
Rust workspace 49 crates, about 397,000 lines (261,000 source, 136,000 tests)
Web UI about 65,000 lines of TypeScript (React, Vite, Playwright)
Test functions 4,613 (static count)
Gate classification corpus 483 cases
Gate bypass corpus 987 cases (must-gate bypasses plus read-only negative controls)
Conformance fixtures 14 (leases, fleet reads, handshake, claims)
Known-failure baseline empty

How we test.

  • Property tests on the gate, the lease rules and secret sealing.
  • Fuzzing of the gate classifier, the usage parsers, and the agent-identity helper's arguments (a million iterations by default). A fuzz crash becomes a regression test before it is fixed.
  • New review findings become failing corpus cases first.
  • Exhaustive matrices, such as every work state against every command.
  • No wall clock in tests. Tests use a fake clock, and a guard test enforces it. Test retries are set to zero in every profile.
  • Conformance as a contract. Fixtures are protobuf JSON; a run that skips a required fixture fails.
  • CI runs on self-hosted Linux arm64 runners only, the main suite runs in about a minute, and a hung test is killed at 300 seconds. macOS coverage runs on maintainers' machines. A public-text guard runs on every pull request.

Choosing the terminal model, measured. Before writing the screen reader we replayed five real recorded Claude Code sessions plus synthetic Codex and stress recordings through three Rust terminal emulators, comparing each against a headless reference emulator. Two identical runs on an M-series Mac:

Terminal crate Text mismatches Unexplained modes Failed checkpoints Throughput (MB/s CPU) Memory per 200×50 screen + 10k scrollback Result
alacritty_terminal 0.26 0 0 0 162.8 49,922 KiB chosen
avt 0.18 0 20 4 81.9 32,194 KiB rejected: no bracketed paste
vt100 0.16.2 144 0 0 85.4 63,328 KiB rejected: unmaintained

The same spike showed that 30 seats with 10,000 lines of scrollback each would cost about 1.5 GB, which is why the default scrollback is 2,000 lines.

What we have not measured yet. Gate decision latency, daemon RPC latency (the histograms exist; we have not published numbers), throughput or merged pull requests per day attributable to the factory, and cost per merged pull request. The last one is the outcome metric we care most about, and the usage ledger exists to feed it.

15. Running our own work on it

On 5 October 2026 we moved our own engineering crew off the earlier platform and onto Shiftyard in one hard switch, with no parallel run. The import moved the open work across and checked that no authority records (grants, approvals) travelled with it. The go/no-go list had six blockers (per-seat accounts, MCP for seats, guardrail parity, Codex wake registration, disk and session resume), and some were repaired in flight.

Since then Shiftyard has largely been built by agent seats running on Shiftyard: the repository has 366 merged pull requests, and the busiest day after the switch had 356 commits. The same mechanisms described above (the gate, the land train, verdicts bound to head commits) govern the agents that are writing the code.

16. Roadmap and Yardcrew

A magenta voyage route joins filled buoys dated Sep 25, Oct 5, and Oct 6, hollow buoys dated Oct 12 and Oct 15, and a dashed continuation to an island labelled Yardcrew.Completed and planned milestones chart the voyage toward Yardcrew.

A 2026 milestone timeline marks completed work from September 25 through October 6, planned October memory migration and retirement, next items in review, and later multi-host, Linux sandbox, and public release work, with a separate indicative Yardcrew lane.The roadmap separates delivered foundations, reviewed next steps, and planned horizons.

Multi-host. A host-to-host link protocol, a fleet read service and conformance fixtures exist; the coordinator that implements them does not. One host acting as the coordinator for a person's other hosts is planned.

Yardcrew is the planned organisation layer on top of Shiftyard. The design installs it in the customer's own AWS account, with yards always dialling out and the control plane never connecting into a yard. The first release is scoped to enrolment, fleet visibility, customer-signed policy (Ed25519 signatures over canonical documents, a host-pinned trust root, effective permissions as the intersection of org and host policy, with the built-in floor always preserved) and exact outcome evidence. Shared claims across hosts, strict fleet budgets, a visual flow builder and organisation-wide memory come later. The indicative path to a first partner pilot is 10 to 12 weeks; it is a plan, not a delivery date, and none of Yardcrew is available today.

17. What we do not claim yet

Engineering posts about agents tend to round up. Here is the list of things Shiftyard does not do today:

  • It is not installable, and the public repository is not open yet.
  • Memory, the code index, embeddings and reranking are not running inside the daemon; the building blocks exist and the cutover is planned.
  • No embedding model has been selected, and we have no retrieval-quality number for Shiftyard's own memory yet.
  • Shiftyard trains no models.
  • OpenCode seats, the heat guard and applying the balancer's moves are in review, not running.
  • There is no Linux sandbox and no multi-host coordination.
  • The gate is not a security boundary, and the verify step's isolation is recorded as weak.
  • We have measured no productivity, cost or quality improvement, and we are not implying one.

Get involved

Shiftyard is licensed Apache-2.0, and the names f200, Shiftyard and Yardcrew are trademarks of f200: forks are welcome under a different name. The public repository is coming soon, with a Developer Certificate of Origin sign-off for contributions. If you run coding agents in parallel and recognise the failures in section 1, or you want to be an early Yardcrew design partner, book a 30-minute call with the founders.