# Iron Reins

Status: implementation-ready research plan  
Prepared: 2026-08-24  
Scope: enterprise software factories and agents that can build and safely operate enterprise-grade, template-constrained EKS platforms

Current phase: **research only**. Iron Reins is cataloging systems, code, operator evidence, commercial boundaries and composition gaps. It is not installing, authenticating to, purchasing or running candidate products.

**Iron Reins** is the one name for the research system, ontology, benchmark, recurring passes and evaluation tools. It asks whether an AI organization can build, ship, secure, operate, support, govern and evolve a bounded enterprise platform while authority, evidence, interruption and recovery remain under control. “Software factory” is one archetype inside the broader Autonomous Technology Organization ontology. Frontier, open-range and first-light language are narrative motifs, not additional product names.

The first qualification lane is deliberately narrower than “replace IT”: support enterprise-grade, cookie-cutter Amazon EKS platforms generated and evolved from versioned Copier templates. Agents work inside a paved road of approved template versions, input schemas, Terraform/Helm/GitOps modules, policy, observability, rollout and rollback contracts. This limits the action space without pretending that a smaller action space makes diagnosis reliable.

**Liferaft** is the name of the modernization and system-replacement lane inside Iron Reins, not a second overall product. It observes bounded incumbent use, specifies replacement contracts, builds and independently verifies a candidate, shadows without side effects, reconciles, canaries by tenant/workflow, transfers authority and retires the incumbent. Network and OTel evidence reveal behavior but never become business truth by themselves.

## Executive answer

Build this as an evidence graph plus a recurring audit loop, following the strongest parts of the State of AI hedge-fund research machine:

1. Define a bounded ontology before collecting products.
2. Capture immutable sources before writing conclusions.
3. Store atomic, time-bounded claims with negative evidence and disconfirmation queries.
4. Score capability, maturity, adoption, autonomy, and evidence strength separately.
5. Reproduce public benchmarks and run factory-level longitudinal evaluations.
6. Promote findings only when a source reveals operating machinery, measured outcomes, or a human authority boundary.
7. Re-run discovery until a bounded pass produces no material additions, then move to a slower maintenance cadence.

The core mistake to avoid is treating every multi-agent coding demo as a software factory. A factory is a persistent production system that accepts demand, moves work through governed stages, produces verifiable artifacts, and learns from failures. A coding harness, agent team, Slack intake bot, refactoring engine, or SRE agent may be a factory component without being a complete factory.

## Research pack

- [`report-source.md`](report-source.md) — canonical findings and direct conclusions;
- [`greenfield-reference-architecture-2026-08-24.md`](greenfield-reference-architecture-2026-08-24.md) — opinionated AWS/EKS/PostgreSQL reference stack, public price anchors, explicit exclusions and uncovered human/commercial functions;
- [`ontology.yaml`](ontology.yaml) — versioned evidence/system/lifecycle/language/model schema;
- [`repository-index.md`](repository-index.md) — canonical, adjacent and private repository boundaries;
- [`commercial-boundaries.md`](commercial-boundaries.md) — stable OSS versus hosted/enterprise gates, deployment/control-plane ownership, dated prices and disaggregated costs;
- [`passes/2026-08-23-pass-001.md`](passes/2026-08-23-pass-001.md) — first bounded discovery/topology pass, contradictions and query mutations;
- [`passes/2026-08-23-pass-002.md`](passes/2026-08-23-pass-002.md) — OSS composition, whole-IT models, operator signals and second-pass discoveries;
- [`passes/2026-08-23-pass-003.md`](passes/2026-08-23-pass-003.md) — Liferaft replacement assurance, OTel/open-observability limits, controlled package ingress and independent build verification;
- [`passes/2026-08-23-pass-004.md`](passes/2026-08-23-pass-004.md) — independent MCP/connector fidelity, proven system-of-record answers, effective-admin lineage and reconciled access auditing;
- [`passes/2026-08-23-pass-005.md`](passes/2026-08-23-pass-005.md) — Microsoft estate operations, cross-suite lifecycle transactions, vendor-change resilience and business-user outcome proof;
- [`passes/2026-08-23-pass-006.md`](passes/2026-08-23-pass-006.md) — bitemporal enterprise facts, transactional outbox/CDC, correction/restatement and honest legacy temporal reconstruction;
- [`passes/2026-08-23-pass-007.md`](passes/2026-08-23-pass-007.md) — call provenance, derived-field lineage, purpose-specific fitness, actual data use and correction impact over decisions/actions;
- [`passes/2026-08-23-pass-008.md`](passes/2026-08-23-pass-008.md) — FinOps, SaaS/COTS/license operations, OSS component boundaries, vendor-collector limits and notebook/TUI research promotion;
- [`passes/2026-08-23-pass-009.md`](passes/2026-08-23-pass-009.md) — OpenAI Kepler/OpenMetadata claim reconciliation, six-layer context, memory and permission-preserving retrieval;
- [`passes/2026-08-23-pass-010.md`](passes/2026-08-23-pass-010.md) — OpenLineage stable-source audit, archived OpenBytes classification and the lineage-assertion supply chain;
- [`passes/2026-08-23-pass-011.md`](passes/2026-08-23-pass-011.md) — data-quality objectives, observations, incidents, repairs, waivers, edition boundaries and full Day-2 closure;
- [`passes/2026-08-24-pass-013.md`](passes/2026-08-24-pass-013.md) — NVIDIA's full public AI-factory stack, AutoML survivors, DSX/NVSentinel code findings and corpus-wide reconciliation;
- [`passes/2026-08-24-recent-turns-coverage-ledger.md`](passes/2026-08-24-recent-turns-coverage-ledger.md) — explicit twenty-thread coverage audit across ontology, dossiers, repositories, commercial boundaries, benchmarks, evidence and queue;
- [`operating-model/whole-it-reference-crosswalk.md`](operating-model/whole-it-reference-crosswalk.md) — IT4IT-backed whole-lifecycle crosswalk and federated authority model;
- [`operating-model/minimum-coherent-technology-organization.md`](operating-model/minimum-coherent-technology-organization.md) — minimum organs, production closure, complexity budget, agent-operability, Crossplane, security overlap, continuous control evidence and financial-services applicability;
- [`operating-model/enterprise-function-control-map.md`](operating-model/enterprise-function-control-map.md) — all technology-department functions, authoritative objects, interactions, human authority and continuous SOC 2 evidence contribution;
- [`operating-model/liferaft-replacement-lane.md`](operating-model/liferaft-replacement-lane.md) — authorized contract mining, shadow/dual-run/canary replacement, authority transfer, rollback and incumbent retirement;
- [`ecosystems/oss-composable-it-stack.md`](ecosystems/oss-composable-it-stack.md) — OSS component map, paid boundaries, deficiencies and assembly hypothesis;
- [`ecosystems/control-plane-substitutability.md`](ecosystems/control-plane-substitutability.md) — OpenTofu/Terraform, Crossplane, Pulumi/CDK, GitOps, portals, operators, secrets and policy by authority, agent-operability, complexity and failure seams;
- [`ecosystems/continuity-and-recovery-organ.md`](ecosystems/continuity-and-recovery-organ.md) — Kubernetes, database, object, SaaS, identity, key, control-plane and audit recovery composition plus the Recovery Gauntlet;
- [`ecosystems/enterprise-data-lifecycle.md`](ecosystems/enterprise-data-lifecycle.md) — D0–D12 data-workflow assessment, factory boundaries, data-plane composition, late-data, database, SQL and vendor-management protocols;
- [`ecosystems/bitemporal-information-and-cdc.md`](ecosystems/bitemporal-information-and-cdc.md) — greenfield two-time truth, CDC/outbox contracts, call and field provenance, purpose-specific use, correction lineage and graded legacy reconstruction;
- [`ecosystems/finops-commercial-assets-and-research-promotion.md`](ecosystems/finops-commercial-assets-and-research-promotion.md) — reconciled commercial estate, OSS FinOps/SaaS/COTS/license components, agent gaps and notebook-to-production path;
- [`ecosystems/data-native-agentic-factories.md`](ecosystems/data-native-agentic-factories.md) — dbt Wizard, Datus, DataSQRL, Kunumi, Altimate, Snowflake, Databricks and Astronomer tools, models, languages, editions, evidence and deficiencies;
- [`ecosystems/kepler-openmetadata-data-agent.md`](ecosystems/kepler-openmetadata-data-agent.md) — Kepler's six-layer data-agent architecture, corrected adoption/scale/latency claims, OpenMetadata stable-source audit and OSS/commercial boundary;
- [`ecosystems/open-lineage-and-data-context-prior-art.md`](ecosystems/open-lineage-and-data-context-prior-art.md) — OpenLineage/OpenBytes/OpenMetadata distinctions, metadata and provenance prior art, OSS comparator queue and assembly rules;
- [`ecosystems/data-quality-operating-system.md`](ecosystems/data-quality-operating-system.md) — quality objectives, observations, incidents, repairs, waivers, OSS/source-available tool boundaries and Day-2 closure;
- [`ecosystems/automl-and-ml-data-quality-survivors.md`](ecosystems/automl-and-ml-data-quality-survivors.md) — living, maintenance and legacy AutoML/search/DQ systems, current license changes and a bounded composable stack;
- [`ecosystems/nvidia-software-and-model-factory.md`](ecosystems/nvidia-software-and-model-factory.md) — NVIDIA's kernels-to-DSX stack, model/data/training/evaluation/inference/agent/Day-2 layers, open/paid boundary and future signals;
- [`evidence/nvidia-github-organization-inventory.md`](evidence/nvidia-github-organization-inventory.md) — dated census of NVIDIA's verified public GitHub organizations with depth tiers and archive/support boundaries;
- [`ecosystems/recent-enterprise-platforms-and-control-surfaces.md`](ecosystems/recent-enterprise-platforms-and-control-surfaces.md) — Netlify Agent Runners, Graphwise, Sentra DSPM and SQL Sentry placed in their correct enterprise organs;
- [`ecosystems/open-observability-and-contract-mining.md`](ecosystems/open-observability-and-contract-mining.md) — OTel/open-source signal plane, eBPF discovery, coverage contracts, OpenSRE boundary and agent-ready evidence;
- [`ecosystems/software-supply-chain-and-build-verification.md`](ecosystems/software-supply-chain-and-build-verification.md) — controlled package ingress, isolated builds, rootless-versus-privileged boundaries and independent GitLab verification;
- [`ecosystems/mcp-and-system-of-record-connectors.md`](ecosystems/mcp-and-system-of-record-connectors.md) — independent MCP/connector assessment, source-answer provenance, delegated identity, administrative authority and access auditing;
- [`ecosystems/microsoft-enterprise-operations-organ.md`](ecosystems/microsoft-enterprise-operations-organ.md) — governed operation of hybrid AD, Entra, Microsoft 365, Intune/Windows, Purview, Power Platform, Azure, licenses and business-user lifecycle;
- [`ecosystems/pass-005-microsoft-specialist-dossiers.md`](ecosystems/pass-005-microsoft-specialist-dossiers.md) — independent product-level reviews of tenant configuration, hybrid identity/recovery, endpoint/application operations and M365 migration/protection;
- [`ecosystems/pass-002-oss-candidate-dossiers.md`](ecosystems/pass-002-oss-candidate-dossiers.md) — tools, languages, models, topology, communication, authority and gaps for new OSS candidates;
- [`methods/operator-signal-smoke-test.md`](methods/operator-signal-smoke-test.md) — mandatory X/Reddit/GitHub/support-forum prequalification lane;
- [`methods/github-huggingface-kaggle-source-lanes.md`](methods/github-huggingface-kaggle-source-lanes.md) — code, issue, model, dataset, notebook and competition discovery with platform-specific evidence gates;
- [`people/builders.md`](people/builders.md) — publicly attributable builders and interview leads;
- [`benchmarks/whole-sdlc-benchmark-stack.md`](benchmarks/whole-sdlc-benchmark-stack.md) — whole-lifecycle benchmark portfolio and 30/90-day protocol;
- [`benchmarks/enterprise-data-lifecycle.md`](benchmarks/enterprise-data-lifecycle.md) — temporal-correctness, migration, performance, cost, lock-in and Day-2 data benchmark;
- [`benchmarks/bitemporal-enterprise-information.md`](benchmarks/bitemporal-enterprise-information.md) — greenfield and legacy valid-time/recorded-time, snapshot/CDC, restatement, uncertainty, privacy and recovery fixture;
- [`benchmarks/commercial-operations-and-research-promotion.md`](benchmarks/commercial-operations-and-research-promotion.md) — cross-source commercial inventory, agreements, entitlements, FinOps action and research-artifact production qualification;
- [`benchmarks/control-plane-complexity-and-production-assurance.md`](benchmarks/control-plane-complexity-and-production-assurance.md) — substitutability, agent-operability, security coverage, recovery, production scope and continuous-control evidence fixture;
- [`benchmarks/liferaft-system-replacement.md`](benchmarks/liferaft-system-replacement.md) — whole-system discovery, contract, shadow, reconciliation, canary, Day-2, recovery and retirement benchmark;
- [`benchmarks/mcp-system-of-record-fidelity.md`](benchmarks/mcp-system-of-record-fidelity.md) — sealed connector-versus-native tests for completeness, correctness, freshness, authorization, provenance and who-accessed-what evidence;
- [`benchmarks/data-agent-context-and-permission-fidelity.md`](benchmarks/data-agent-context-and-permission-fidelity.md) — semantic retrieval, query correctness, memory, per-layer permission, provenance and latency/cost qualification;
- [`benchmarks/lineage-context-and-provenance-fidelity.md`](benchmarks/lineage-context-and-provenance-fidelity.md) — event/edge completeness, field provenance, permission safety, conflict, replay, recovery and agent-answer qualification;
- [`benchmarks/data-quality-operations-and-repair.md`](benchmarks/data-quality-operations-and-repair.md) — fault precision/recall, false alerts, causal diagnosis, blast radius, repair, restatement, reconciliation and quality-system Day-2 qualification;
- [`benchmarks/nvidia-and-model-factory-lifecycle.md`](benchmarks/nvidia-and-model-factory-lifecycle.md) — artifact rights, curation, training, evaluation, promotion, serving, DSX event semantics, GPU Day 2 and inference portability;
- [`benchmarks/microsoft-estate-autonomy.md`](benchmarks/microsoft-estate-autonomy.md) — hybrid Microsoft estate lifecycle, administration, partial-failure, vendor-change, user-journey and audit qualification;
- [`benchmarks/eks-template-operations.md`](benchmarks/eks-template-operations.md) — primary template-constrained EKS build, upgrade, RCA and remediation qualification lane;
- [`factorybench/mlflow-ec2-to-multitenant-kubernetes.md`](factorybench/mlflow-ec2-to-multitenant-kubernetes.md) — per-ecosystem comparative stress test, not a migration plan;
- [`infrastructure/factory-qualification-lab.md`](infrastructure/factory-qualification-lab.md) — Cloudflare OS/Sandbox, Modal and real-cloud/Windows qualification tiers;
- [`infrastructure/disposable-enterprise-beta-lab.md`](infrastructure/disposable-enterprise-beta-lab.md) — future-only qualification protocol; execution is outside the current phase;
- [`ecosystems/`](ecosystems/) — source-backed dossiers for internal, open, local, Day-2 and Microsoft ecosystems;
- [`ecosystems/whole-it-autonomy.md`](ecosystems/whole-it-autonomy.md) — contenders explicitly pursuing autonomous or whole-function enterprise IT;
- [`research-queue.md`](research-queue.md) — reproduction, expansion and evidence gaps.

## Decision contract

Iron Reins must answer, with dated evidence:

- What problem is each system trying to solve?
- What enters the system, and what exits?
- Which agent and subagent roles exist, which are real execution units versus prompt personas, and who may spawn whom?
- How do agents communicate, claim work, exchange artifacts, acknowledge or reject handoffs, resolve conflicts, and recover from lost or duplicated messages?
- What context and memory can each role read or write, and how are staleness, poisoning, compaction, isolation and replay handled?
- Which software-lifecycle stages are actually automated?
- How long can it operate without human attention?
- Where are humans required, and what authority do they retain?
- Does it deploy, operate, diagnose, remediate, and learn after release?
- How does it prevent functional defects, security failures, architectural erosion, duplication, and review overload?
- What evidence shows adoption: source code, dogfooding, public traces, production throughput, customer use, or independent evaluation?
- Has the complete system, not merely its underlying model, been evaluated on a relevant benchmark?
- Can it improve a legacy codebase, or does it merely reproduce the codebase's existing mud?
- What fails, what cannot be inferred, and what evidence would change the conclusion?
- Which technology-department functions can it observe, recommend for, execute within, own as a gated workflow, own to an SLO, or govern under policy and budget?
- What human staffing, expertise, review, escalation, on-call and vendor support remain after deployment?
- If an OSS project has paid editions, exactly which capabilities, limits and operational responsibilities move behind the commercial boundary?
- What does each edition cost under explicit team-size, usage, infrastructure, support and contract-term assumptions?
- Are software, model inference, agent execution, target infrastructure, support, implementation and residual-human costs reported separately rather than hidden inside one TCO number?
- Which signals are emitted, sampled, transformed, dropped and retained; how is OTel semantic and population completeness measured rather than assumed?
- Which incumbent behaviors were directly observed, which were supplied by contracts or humans, and which remain invisible or unknown?
- Can it shadow without side effects, reconcile state and business invariants, canary a complete authority domain, roll back and retire the incumbent?
- Can an agent or build bypass approved package ingress, poison shared caches, alter protected checks, inspect hidden tests or publish a trusted artifact?
- Is each MCP/API/CLI connector assessed independently from its underlying system, using the exact connector and source-API versions?
- Can every system-of-record answer identify its principal/delegation chain, tenant/snapshot, exact query/pages, source object IDs and versions, omissions, freshness and native audit evidence?
- Can the system answer who held effective administrative authority at a requested time, why they held it, who approved it and whether it was revoked or recertified?
- Can it reconcile who accessed, exported, changed, deleted, approved or administered what across agent, host, gateway, connector and source-system audit records?
- Can it reconcile every SaaS/COTS/cloud/license asset from agreement and invoice through assignment, meaningful use, renewal, safe reclamation, realized optimization and exit, without trusting one vendor collector?
- Does a research artifact have a tested path from notebook or interactive/TUI work into a package, typed machine interface, Dagster-owned execution, deployment, observability, recovery, support and field-level provenance?

## Inclusion boundary

Include a system when it does at least one of the following and exposes enough machinery to assess:

- moves an issue, specification, alert, or change request through two or more software-delivery stages;
- coordinates persistent or parallel coding agents with state, isolation, recovery, and verification;
- encodes a repeatable software process as workflows, contracts, policies, skills, or agent roles;
- performs large-scale deterministic or agentic modernization across repositories;
- operates deployed software through monitoring, deployment verification, incident investigation, remediation, rollback, or prevention;
- closes a learning loop by turning failures, reviews, or incidents into evals, rules, memory, or workflow changes.

Exclude or classify as adjacent:

- a standalone coding model or IDE chat mode;
- a generic multi-agent framework with no software-production use;
- a demo that produces code but has no execution or verification evidence;
- a CI product with no agentic planning or adaptation;
- marketing claims without a named workflow, artifact, control, or outcome;
- benchmark results for a base model presented as proof of the factory around it.

## Ontology

### Factory archetypes

One system may have several archetypes, but one must be marked primary.

| Code | Archetype | Defining output |
|---|---|---|
| `F1` | Interactive coding harness | Human-steered code changes in a working repository |
| `F2` | Agent-team/orchestration substrate | Coordinated agent sessions, tasks, messages, and workspaces |
| `F3` | Workflow-as-code factory | Versioned graph or pipeline that executes agents, commands, and gates |
| `F4` | Issue-to-PR factory | Reviewed branch or pull request from issue/spec intake |
| `F5` | Full-SDLC factory | Intake through validation, release, documentation, and monitoring |
| `F6` | Self-improving factory | Failures or reviews become evals, policies, context, or workflow changes |
| `F7` | Specification/contract factory | Executable requirements and boundary contracts govern generation |
| `F8` | Modernization factory | Behavior-preserving migration, refactoring, or decomposition at scale |
| `F9` | Verification factory | Reproduction, testing, security, live-environment evidence, and review |
| `F10` | Release/supply-chain factory | Build, package, sign, provenance, publish, canary, and rollback |
| `F11` | Day-2/SRE factory | Observe, investigate, remediate, recover, and prevent incidents |
| `F12` | Work-intake/ambient agent | Persistent organizational interface that routes work into other systems |

### Lifecycle coverage

Assess every stage as `none`, `assistive`, `delegated`, `autonomous-gated`, or `autonomous-dark`, and attach evidence.

| Code | Stage | Required evidence examples |
|---|---|---|
| `L0` | Demand sensing | Issue, alert, telemetry, support, security, or dependency signal |
| `L1` | Intake and normalization | GitHub, Jira, Linear, Slack, Sentry, API, or contract ingestion |
| `L2` | Triage and prioritization | Classification, duplicate detection, risk/value ranking |
| `L3` | Discovery and specification | Reproduction, business rules, acceptance criteria, architecture fit |
| `L4` | Planning and decomposition | Typed plan, DAG, tasks, dependencies, budget, ownership |
| `L5` | Workspace and context | Checkout/worktree, sandbox, credentials, network policy, memory |
| `L6` | Implementation | Code, tests, configuration, infrastructure, docs |
| `L7` | Deterministic verification | Build, tests, types, lint, security, mutation, contract checks |
| `L8` | Independent review | Separate context/model, risk assessment, evidence chain |
| `L9` | Merge and release | Approval, merge, version, artifact, SBOM, signing, provenance |
| `L10` | Deploy and verify | Preview, staging, production, canary, health and behavior checks |
| `L11` | Operate and respond | Telemetry, alert triage, RCA, remediation, rollback, communication |
| `L12` | Learn and improve | Failure taxonomy, memory, eval generation, policy/workflow updates |

### Autonomy levels

| Level | Meaning |
|---|---|
| `A0` | Suggests; human executes every material action |
| `A1` | Executes one bounded task under active supervision |
| `A2` | Runs a repeatable multi-stage workflow with human gates |
| `A3` | Runs unattended to a reviewed draft PR or equivalent artifact |
| `A4` | Can merge, release, or deploy within policy and approval boundaries |
| `A5` | Detects, changes, deploys, operates, and learns in a closed loop |

Autonomy is recorded per lifecycle stage. A single global autonomy label is prohibited.

### Maturity levels

| Level | Meaning |
|---|---|
| `M0` | Concept or marketing claim |
| `M1` | Runnable prototype or source-backed demo |
| `M2` | Dogfooded on the builder's own repositories |
| `M3` | External pilot or named customer workflow |
| `M4` | Repeated production use with dated throughput/reliability evidence |
| `M5` | Reproducible independent evaluation plus sustained multi-team adoption |

### Evidence strength

| Tier | Evidence |
|---|---|
| `E0` | Unverified assertion, post, or search lead |
| `E1` | Official description, documentation, source, or release artifact |
| `E2` | Reproducible local run, trace, test fixture, or benchmark submission |
| `E3` | Dated dogfood telemetry, public PR history, or production throughput |
| `E4` | Independent benchmark, controlled comparison, or customer-confirmed result |
| `E5` | Audited longitudinal evidence across organizations and failure conditions |

Do not collapse maturity and evidence. A closed enterprise product may be mature with weak public evidence; an open-source prototype may have excellent reproducibility but little adoption.

## The Day-2 test

A system only receives Day-2 credit for capabilities demonstrated after deployment:

- watches application and infrastructure health;
- correlates telemetry, deploys, commits, configuration, and topology;
- verifies a release in a realistic environment;
- detects regressions and opens a governed work item;
- develops and tests competing root-cause hypotheses;
- proposes or executes a bounded remediation;
- performs or recommends rollback;
- updates responders and stakeholders;
- generates a postmortem and tracks prevention work;
- turns an incident into a regression test, runbook, guardrail, or factory change.

Merely running the factory server, retrying an agent, or monitoring token use is **factory operations**, not **application Day-2 operations**. Record both, but never conflate them.

### Root-cause discipline

RCA is not free-form explanation generation. Iron Reins scores four outcomes independently:

1. **localization** — where the failure manifests;
2. **causal diagnosis** — the smallest cause that explains the observations and counterfactuals;
3. **remediation selection** — the least invasive approved change that addresses that cause;
4. **verified recovery** — restoration, no unacceptable collateral regression, and no recurrence during the observation window.

A correct fix with an incorrect causal explanation does not receive RCA credit. A symptom-suppressing restart, scaling change, timeout increase or rollback is mitigation, not a root-cause fix. An inaccurate diagnosis that causes a mutation is a safety failure even if the service later becomes healthy.

For the EKS lane, production mutation is prohibited until the system records competing hypotheses, evidence for and against each, information still missing, confidence/calibration, and a discriminating test. The mutation must target an allowlisted Copier input, template/module version, GitOps object or runbook action; direct imperative cluster edits are emergency mitigations that must be reconciled back to desired state.

## The slop-control model

Assess slop across four layers.

### 1. Functional correctness

- characterization, unit, integration, end-to-end, and property tests;
- test adequacy and mutation score, not only green status;
- reproducible bug probes and hidden or adversarial tests;
- live behavior checks where mocks would hide integration failures.

### 2. Structural integrity

- complexity and change concentration;
- duplication and copy-paste growth;
- module boundaries and dependency direction;
- contract/API compatibility;
- dead code, abstraction count, and diff surface;
- architectural fitness functions and forbidden dependencies.

### 3. Operational and security integrity

- secret isolation and least privilege;
- sandbox and network containment;
- dependency, license, vulnerability, and supply-chain checks;
- performance budgets and resource ceilings;
- canary, rollback, and blast-radius policies;
- provenance for prompts, tools, models, commits, artifacts, and approvals.

### 4. Learning integrity

- each rejected or reverted change receives a failure class;
- failures become counterexamples and regression fixtures;
- independent review is separated from implementation context;
- soft LLM judgments are calibrated against deterministic and human labels;
- prompts, skills, policies, and evals are versioned and linked to outcomes;
- Goodhart-resistant metrics retain denominators, abstentions, and negative outcomes.

SlopCodeBench is an important external signal because it measures iterative structural erosion and verbosity, but it is not sufficient. Its May 2026 revision reports that no evaluated agent solved a complete problem end-to-end, and that structural erosion or verbosity worsened in most trajectories. Iron Reins should therefore run a factory-level variant in which the same evolving product passes through the factory's full verification and learning loop, not only its underlying coding CLI. Source: [SlopCodeBench paper](https://arxiv.org/abs/2603.24755) and [leaderboard](https://www.scbench.ai/).

## Human authority map

For every system, identify whether a human is required to:

- choose work;
- define or approve the specification;
- resolve product ambiguity;
- approve the plan;
- grant credentials or permissions;
- judge architecture and risk;
- review code or evidence;
- merge;
- release or deploy;
- authorize production remediation or rollback;
- declare an incident resolved;
- accept a factory rule or eval change.

Record both the documented policy and observed practice. “Human in the loop” is too vague to score.

## Adoption and outcome model

Prefer behavioral measures over registrations, stars, or vendor customer logos.

### Throughput

- accepted tasks per week;
- percentage of merged PRs authored or materially completed by the factory;
- percentage of issues closed with factory evidence;
- median lead time and review time;
- parallelism, queue age, and work-in-progress.

### Quality and reliability

- first-pass acceptance and revision count;
- defect escape and revert rate;
- change-failure rate;
- security findings per accepted change;
- architecture/complexity/duplication deltas;
- deployment success, rollback, and MTTR.

### Economics and human load

- model, compute, sandbox, and tool cost per accepted change;
- cost per failed, reverted, or abandoned run;
- human minutes per accepted change and per incident;
- reviewer queue depth and review-compression ratio;
- cost of operating and improving the factory itself.

### Adoption depth

- weekly active users and active repositories;
- retention after 4, 12, and 26 weeks;
- number of teams using standard workflows versus one champion;
- percentage of runs using production integrations and real test environments;
- concentration: whether one maintainer accounts for most accepted output.

## Benchmark program

No public benchmark currently answers whether an enterprise software factory remains useful, secure, and maintainable over months. Use a portfolio.

| Question | External benchmark or evidence |
|---|---|
| How long a human-calibrated task can it complete reliably? | METR Time Horizon 1.1 and HCAST |
| Can it resolve bounded repository issues? | SWE-bench Verified/Multilingual and SWE-bench Pro |
| Can it sustain release-scale or chained evolution? | SWE-EVO, RoadmapBench, SWE-Chain and EvoClaw |
| Does its code erode under repeated extension? | SlopCodeBench plus architecture and maintainability fitness tests |
| Can it span build, configuration, monitoring, issue repair and tests? | DevOps-Gym's 700+ Java/Go tasks and chained flows |
| Can an independent agent review consequential defects? | c-CRAB code-review benchmark plus blinded maintainer review |
| Is a functional change also secure? | SecureAgentBench exploit/static-analysis gates and organization threat fixtures |
| Can it diagnose and safely resolve production failures? | AIOpsLab, ITBench, SREGym, OpenSRE and live incident replays |
| Can it modernize without semantic loss? | Differential behavior suites, golden-master replays, contract and data reconciliation fixtures |
| Does it help over real calendar time? | 30-day shadow and 90-day canary trials using DORA's five metrics plus quality, cost and human-load telemetry |

Every run must capture factory version, model/provider, prompts/skills, repository snapshot, sandbox image, tools, budget, wall time, retries, human interventions, patch, traces, test logs, and cost. Custom scaffolds and standard scaffolds are separate leaderboards.

See [`benchmarks/whole-sdlc-benchmark-stack.md`](benchmarks/whole-sdlc-benchmark-stack.md) for task counts, languages, limitations, contamination controls, conjunctive pass criteria and the longitudinal protocol.

### FactoryBench: internal longitudinal suite

Create a StateBench-like suite with five tracks:

1. **Brownfield evolution** — 12 checkpoint sequences over real service code, with hidden behavior and architecture tests.
2. **Big-ball-of-mud rescue** — unstable legacy applications requiring characterization, seams, modularization, and staged extraction.
3. **Secure delivery** — issue to signed artifact, staging deployment, smoke test, canary decision, and rollback.
4. **Template-constrained EKS operations** — alert to evidence-grounded RCA, bounded desired-state remediation, verification, rollback/reconciliation, and postmortem.
5. **Factory learning** — repeat earlier failure classes after the system receives corrections; measure recurrence and collateral regressions.

Score correctness, maintainability delta, security, operations, human minutes, cost, and evidence completeness separately. A single aggregate score can be published only with all component scores and weights.

## Legacy modernization: from mud to an agent-ready system

Do not begin by asking an agent to “clean up” a large codebase. Use a gated migration factory.

### Stage 0: freeze the evidence, not the business

- inventory repositories, builds, deployments, databases, queues, APIs, jobs, owners, and runtime dependencies;
- establish a reproducible build and executable environment;
- capture production traffic shapes, critical reports, and operational runbooks;
- identify legal, regulatory, latency, availability, and data-retention constraints.

### Stage 1: recover behavior and contracts

- generate characterization and golden-master tests around observable behavior;
- record data schemas, rounding, ordering, time, locale, and failure semantics;
- extract business rules with source spans and domain-expert confirmation;
- map callers, dependencies, change coupling, incidents, and hot spots;
- treat reverse-engineered specifications as hypotheses until tests and experts agree.

Reversa is a useful early candidate for reverse-documentation contracts; AWS Transform and IBM watsonx are enterprise candidates for mainframe and language migration; OpenRewrite/Moderne provides deterministic, type-aware mass refactoring. These should be tested as components, not assumed to solve the entire modernization program. Sources: [Reversa](https://github.com/sandeco/reversa), [AWS Transform](https://aws.amazon.com/transform/), [OpenRewrite](https://github.com/openrewrite/rewrite), and [IBM watsonx Code Assistant for Z](https://www.ibm.com/docs/en/watsonx/watsonx-code-assistant-4z/).

### Stage 2: create seams and fitness functions

- place contracts around modules, APIs, events, databases, and batch boundaries;
- isolate volatile infrastructure behind ports/adapters where that matches observed behavior;
- add dependency, architecture, performance, and security fitness tests;
- build a semantic diff harness that compares old and new behavior over representative and adversarial inputs.

### Stage 3: transform in bounded slices

- prioritize high-change/high-risk modules with observable boundaries;
- use deterministic recipes before generative edits;
- run one vertical slice through analyze, specify, transform, reconcile, deploy, and observe;
- retain parallel run, shadow traffic, feature flags, and rollback until operational evidence is sufficient;
- require domain-owner approval for behavior changes, even when tests pass.

### Stage 4: prevent remudding

- make contracts and fitness functions required merge gates;
- track coupling, duplication, complexity, change concentration, and service ownership;
- convert incidents and review findings into regression fixtures;
- budget architecture work as a factory lane rather than expecting incidental cleanup from coding agents.

## Initial named-system assessment

This is a seed, not a final ranking. “Unknown” means public evidence is insufficient, not that the capability does not exist.

| System | Best current classification | Public maturity/evidence | Day-2 application operations | Slop/quality controls | Human boundary | Benchmark/adoption signal |
|---|---|---|---|---|---|---|
| Stripe Minions | `F4/F6` proprietary internal PR factory | `M4/E2–E3`: deterministic stages around a coding agent, isolated devboxes and curated internal tools | No public closed production-operations loop | conditional repository rules, targeted lint, selective CI, bounded repair and human review | humans decide whether to open, review, iterate or take over the PR | Stripe reports more than 1,000 wholly Minion-produced, human-reviewed merged PRs per week; no public factory benchmark |
| Ramp Inspect | `F4/F5/F11` proprietary internal background factory | `M4/E2–E3` coding; `M3` bounded telemetry-triggered repair | **Yes, bounded:** generated monitor → alert → sandbox reproduction → fix proposal → human merge | real telemetry, sandboxes, tests and specialist model stages; humans author trusted instrumentation | humans merge and retain observability/production authority | Ramp reports about 30% of merged frontend/backend PRs and 40 real bugs in the first week of its monitor trial; noisy thresholds/duplicates documented |
| Spotify Honk/Fleetshift | `F4/F8` fleet-migration factory | `M4/E2–E3`: 1,500+ reported merged Honk PRs and hundreds of users | No live incident-remediation loop shown | repository-sensitive checks, stop hooks, restricted tools and an independent judge | engineers define migrations, supply missing context and merge | strongest published migration/verifier evidence; no public source or factory-level benchmark |
| Block Goose/G2 | `F2` open substrate plus private workflow layer | `M4` as agent substrate; `M1–M2` as a factory without added systems | technically extensible, but no measured closed loop | permissions/extensions are configurable; isolation and independent verification are external | user authority by default unless governed externally | 15+ providers and broad reported use; no target-language matrix or complete-factory benchmark |
| Open Agents (Vercel Labs) | `F1/F2` cloud-agent reference platform | `M1–M2/E1`: substantial open reference implementation, not a turnkey enterprise factory | No continuous application-operations loop | per-command approvals, path controls, isolated resumable sandboxes, repository-native checks when invoked | dangerous/unknown shell work requests approval; PR creation is optional | one canonical repo found; no factory benchmark or external adoption evidence yet; [repo](https://github.com/vercel-labs/open-agents) |
| Warp Factories / Oz | `F3/F5/F6` control plane | early access; Warp's own repository is the main public dogfood case | Factory design includes monitor-to-issue loops; broad production remediation is not yet demonstrated | factory-as-code, human checkpoints, review/verification agents, evals, canaries/rollback of factory definitions | teams choose checkpoints and can steer/handoff sessions | Warp reports 20–30% of its PRs automated; independent denominator and longitudinal quality audit remain open; [factory guide](https://www.warp.dev/blog/a-guide-to-cloud-software-factories-for-engineering-leaders) |
| Vercel AI SDK Factory | `F4/F6/F9` | `M4/E3`: four weeks of public production reporting | No application-operations loop shown; factory-failure learning is strong | bug probes, isolated sandboxes, independent review, live E2E evidence | humans merge every change | Vercel reports 25–35% of weekly merged PRs authored by the factory and over 75% of July issue closures; [official report](https://vercel.com/blog/building-a-software-factory-for-ai-sdk) |
| Vercel eve Software Factory template | `F4` | `M1–M2/E1`: open template | No | four specialist agents, independent reviewer, sandbox, up to two revisions | stops at draft PR; human marks ready and merges | public source/template, adoption unknown; [template](https://vercel.com/templates/other/eve-software-factory) |
| Fabro | `F3/F9` | `M2/E2`: open, self-hostable, active source | Can orchestrate commands, but no demonstrated application-operations product loop yet | deterministic graphs, sandboxes, git checkpoints, tests/linters/judges, retrospectives | configurable human gates | public SWE-bench Lite harness and scoreboard; no verified SlopCodeBench result found; [repo](https://github.com/fabro-sh/fabro) |
| Tessl Agent | `F6/F7` | `M1–M2/E1`: product and dogfood narrative | No direct application Day-2 proof found | review loops generate eval scenarios; skills/context/evals are portable files | humans approve improvement PRs | external adoption and reproducible factory benchmark remain open; [official explanation](https://tessl.io/podcast/112) |
| Mastra Factory | `F4`, with `F5/F11` reference architecture | `M1/E1`: open alpha | Monitoring and Sentry loops are in the reference design; production proof is open | typed workflows, Zod handoffs, memory, traces, goal judge | configurable; product currently issue-to-PR | alpha, 19 GitHub stars at retrieval, no independent factory benchmark found; [repo](https://github.com/mastra-ai/softwarefactory-template) |
| Uncle Bob SwarmForge | `F2/F7/F9` | `M1–M2/E1`: runnable branch workflows and self-dogfood intent | No | TDD, Gherkin, cleanup, architecture, mutation hardening, QA roles | observable local tmux workflow; human authority retained | no public long-horizon or SlopCodeBench result found; [repo](https://github.com/unclebob/swarm-forge) |
| Agentic Coding Flywheel: NTM + Mail/Beads/UBS/ACFS/DSR and adjacent tools | `F2/F9/F10` ecosystem | component maturity varies; public source and local operational evidence | Strong factory/release operations; application Day-2 depends on attached tools | work ownership, file reservations, bug scans, batch verification, release/SBOM/provenance tooling | operator owns dispatch, acceptance, and risky actions | ecosystem claims 3+ hours autonomous operation and 5.4K stars; factory-level independent benchmark is open; [overview](https://agent-flywheel.com/flywheel) |
| Dana Marble's contract/hexagonal approach | provisional `F7` | `E0–E1`: public writing establishes the problem framing; exact system is not sufficiently documented publicly | Unknown | contract and boundary thesis requires artifact-level verification | Unknown | interview/demo/source needed; [public writing](https://medium.com/@dana.marble/phenotype-evolution-of-applications-produced-by-generative-ai-c81d04a9690d) |
| Atlassian HULA / Kun Chen et al. | `F4` human-in-the-loop | `M3/E2–E3`: deployed internally in Jira and academically described | No | human plan/code refinement, unit testing, functional evaluation | human control at planning and coding stages | paper reports 37.2% on SWE-bench Verified; [paper](https://arxiv.org/abs/2506.11009) |
| Claude Code Agent Teams | `F2` substrate | experimental, disabled by default | No | shared tasks/messages; quality depends on configured roles and tools | interactive lead and permissions | not a factory; known resumption, coordination, shutdown, token, and file-conflict limitations; [docs](https://code.claude.com/docs/en/agent-teams) |
| Claude Tag | `F12` adjacent intake/ambient layer | beta for Team/Enterprise | Can monitor channels and follow deployments, but not an SRE system by itself | skills, instructions, tool identity, credential logging | tags humans when decisions are needed | can turn a thread into a draft PR and run multi-day tasks; adoption claims require separate verification; [official page](https://claude.com/product/tag) |
| Stakpak Agent | `F10/F11` | open source, active, 24/7 autopilot | Yes: deployment watch, infrastructure, Kubernetes, CI/CD, scheduled health work | secret substitution, network guardrails, rulebooks, profiles, sandboxing | policy/approval dependent | no relevant software-evolution benchmark found; [repo](https://github.com/stakpak/agent) |
| HolmesGPT | `F11` | CNCF Sandbox, open source | Yes: investigation, root cause, deployment verification, optional remediation | evidence from connected telemetry and health checks; remediation is separately controlled | configurable human/remediation boundary | adoption and RCA accuracy need scenario-level evidence; [repo](https://github.com/HolmesGPT/holmesgpt) |
| OpenSRE | `F11` research/eval substrate | open and early | Yes, focused on incident investigation and response | scored synthetic RCA plus cloud-backed E2E scenarios | configurable | especially valuable as an evaluation environment, not yet proof of broad adoption; [repo](https://github.com/Tracer-Cloud/opensre) |
| Cloudflare OS | `F1/F5/F12` isolate-based workspace/Gadget platform | open source, new | platform state and approval workflows, not deployed-application SRE | capability Gatekeepers, narrow bindings, outbound controls and queued human approval | humans approve deferred side effects; simulation can diverge before real write | no factory benchmark; Worker-compatible JS/TS scope and isolate constraints make it a separate evaluation cohort |
| Cloudflare Sandbox/Containers | `F1` qualification substrate | preview but substantial public infrastructure | sandbox lifecycle only | per-sandbox VM/filesystem/process/network isolation, custom images, snapshots and egress controls | factory supplies authorization policy | up to 1,500 vCPU/6 TiB/30 TB account ceilings; not itself a software factory |
| AWS DevOps Agent | `F11` managed Day-2 product claim | `M0–M1/E1` on currently retained evidence; not recommended | vendor describes triage, RCA and mitigation; Iron Reins has no acceptable operator-quality proof and received direct negative operator feedback | vendor-described evidence journal and topology are claims to test, not quality evidence | no production authority in Iron Reins; sealed read-only negative-control use only | recommendation withdrawn pending successful sealed RCA/safety trials and corroborated customer outcomes; [docs](https://docs.aws.amazon.com/devopsagent/latest/userguide/production-operations-autonomous-incident-response.html) |
| AWS Transform | `F8` | enterprise managed modernization system | Not post-migration application operations | discovery, planning, transformation, tests, readiness validation; human/domain verification remains critical | modernization teams review plans and outputs | vendor-reported speedups; benchmark-quality semantic equivalence evidence must be audited; [official page](https://aws.amazon.com/transform/) |
| OpenRewrite / Moderne | `F8` deterministic substrate | mature open-source transformation engine plus commercial multi-repo platform | No | type-aware lossless semantic trees and deterministic recipes | humans choose/approve recipes and rollout | strong reproducibility; not a complete agent factory; [repo](https://github.com/openrewrite/rewrite) |

## Research system architecture

Use append-only JSONL as the evidence source of truth and SQLite/DuckDB as rebuildable query projections.

```text
iron-reins/
  AGENTS.md
  README.md
  docs/
    decision-contract.md
    source-policy.md
    scoring-rubric.md
    recurring-audit.md
  ontology/
    software-factory-ontology.yaml
    versions/
  sources/
    web/
    repositories/
    releases/
    benchmarks/
    transcripts/
    interviews/
  evidence/
    records.jsonl
    corrections.jsonl
    queries.jsonl
    decisions.jsonl
  research/
    systems/
    archetypes/
    day-2/
    modernization/
    benchmarks/
  evals/
    factorybench/
    reflect-audit/
  findings/
  wiki/
  site/
  scripts/
  logs/
```

### Canonical records

| Record | Purpose | Required fields |
|---|---|---|
| `SourceArtifact` | Immutable retrieved source | canonical URL/path, hash, retrieval time, published time, source family, access class, authority tier |
| `System` | Product/project/framework identity | canonical name, aliases, owner, archetype, status |
| `OfferingEdition` / `EditionFeatureBoundary` | OSS, hosted and enterprise product boundary | exact edition, entitlement, limit, source and date |
| `PriceObservation` / `CostLedgerObservation` | Dated pricing and TCO components | basis, assumptions, bundled classes, unknowns and source |
| `Person` / `OrganizationRoleInterval` / `SystemContribution` | Public builder provenance | attributable artifact, contribution type, dates, confidence and negative boundary |
| `InterviewLead` | Evidence-oriented outreach queue | rationale, unanswered questions, public contact surface and status |
| `Release` | Time-bounded implementation state | version/tag/commit/date, support status |
| `CapabilityClaim` | Atomic lifecycle or control claim | subject, predicate, stage, scope, time, source spans, confidence, negative boundary |
| `Deployment` | Use in a real organization/repository | adopter, scope, dates, environment, volume, evidence, vendor relationship |
| `HumanGate` | Authority boundary | stage, role, required/optional, action controlled, bypass policy |
| `BenchmarkRun` | Reproducible evaluation | benchmark/version, scaffold, model, system version, config, artifacts, score, cost |
| `Failure` | Observed failure or limitation | class, trigger, consequence, detection, recovery, recurrence |
| `Correction` | Human refinement | prior record, correction type, proposed rule, counterexample test |
| `Query` | Next information-gain action | parent records, source lane, expected result, disconfirmation target, status |
| `Decision` | Assessment or recommendation | rubric version, component scores, evidence coverage, abstentions, date |

## Acquisition lanes

Run these independently so one noisy source family does not dominate:

1. Stable release source audits: repository tags, code, tests, manifests, security model, hidden/experimental flags.
2. Public production traces: authored PRs, issue closures, backports, bot identities, deploy histories, incident examples.
3. Official architecture material: docs, engineering blogs, talks, workshops, transcripts.
4. Independent evidence: papers, benchmarks, reproduction repos, controlled studies, public customer engineering accounts.
5. Adoption: package downloads, releases, contributors, active installations where public, repeated customer use, retention signals.
6. Failure evidence: issues, rejected PRs, incident reports, security advisories, reverts, benchmark trajectories.
7. People/artifact descent: identify builders through attributable commits, designs, reviews, talks, and production ownership; use this for interviews, not automated employment ranking.
8. Commercial-boundary audit: stable OSS tag and license, hosted and enterprise entitlement tables, archived pricing, quote status, support/SLA terms, usage limits and downgrade/export behavior.

For open-source tools, check out the latest stable tag rather than `main`, record the commit, and inspect the implementation before accepting documentation claims. For hosted systems, preserve dated pages and observable public artifacts.

## Recurring control loop

```text
decision contract
  -> source-surface map
  -> parallel discovery/acquisition
  -> immutable artifacts
  -> entity/release resolution
  -> atomic claims + contradictions + negative evidence
  -> per-stage scoring + adoption + human map
  -> benchmark/reproduction queue
  -> human corrections
  -> regression fixtures + query mutations
  -> publication candidates
  -> next pass
```

### Promotion rule

Promote a finding only when it contains at least one:

- named lifecycle stage with a concrete input and output;
- named deterministic, security, operational, or learning control;
- named human authority boundary;
- observable deployment/adoption measure with denominator and date;
- reproducible benchmark configuration and artifacts;
- observed failure, limitation, or counterexample that changes the assessment.

Reject generic autonomy language, vendor logos without a named workflow, model scores attributed to a factory, and production claims that lack dates or scope.

### Bounded audit schedule

- Bootstrap: two passes per day for 14 days.
- Expansion: daily passes until three consecutive completed passes add no material system, capability, contradiction, or benchmark evidence.
- Maintenance: weekly source/release check and monthly benchmark/adoption refresh.
- Event-triggered: new release, benchmark publication, security incident, major customer announcement, acquisition, deprecation, or pricing/model-policy change.
- Hard stop each run: fixed query budget, canonical URL and artifact dedupe, and explicit `no_material_findings` result.

The recurring job may update evidence ledgers and draft assessments. It must not silently change published rankings, delete prior claims, or promote vendor statements without passing source and evidence gates.

## Scoring

Publish a profile, not one magic number. Suggested dimensions, each 0–5:

- lifecycle breadth;
- execution depth;
- verification rigor;
- security and isolation;
- recoverability and durability;
- human-governance clarity;
- Day-2 capability;
- learning-loop quality;
- adoption depth;
- benchmark strength;
- interoperability/portability;
- legacy-code suitability;
- coordination and handoff rigor;
- topology observability and recoverability;
- autonomous technology-organization functional coverage and ownership depth;
- capacity, economics, staffing and organizational resilience.
- OSS completeness, enterprise-feature dependence, price transparency and total cost of ownership.

Beside every score show evidence coverage from 0–100%. A score with less than 40% coverage is `provisional`; less than 20% is `insufficient evidence`. Missing public evidence is not scored as zero capability.

## Deliverables

### Research products

- ontology and glossary;
- exhaustive system registry with aliases and status;
- one evidence-backed card per system and release;
- capability heatmap by lifecycle stage;
- Day-2/SRE landscape;
- legacy-modernization playbook and vendor/component map;
- slop-control comparison;
- human-authority map;
- adoption and throughput ledger;
- OSS-versus-paid feature matrices and dated price observations;
- benchmark crosswalk and reproducibility ledger;
- failure-mode and counterevidence atlas;
- quarterly “state of software factories” synthesis.

### Decision products

- enterprise build/buy/compose decision tree;
- readiness assessment for a target codebase;
- pilot design with human gates and stop conditions;
- reference architecture for a composable factory;
- procurement questions and evidence requests;
- modernization factory blueprint for big-ball-of-mud systems;
- Day-2 operating model and incident authority matrix.

## Execution phases

### Phase 0 — Contract and ontology (week 1)

- approve scope, terminology, inclusion/exclusion rules, evidence tiers, and ethics;
- implement schema validation and canonical IDs;
- create seed roster and source-surface map;
- define materiality and saturation rules.

Exit: ten known systems can be represented without free-text exceptions.

### Phase 1 — Evidence spine (weeks 1–2)

- build source capture, hashing, URL canonicalization, and release resolution;
- add append-only claims, contradictions, corrections, and query derivation;
- create a rebuildable SQLite/DuckDB projection and basic reports.

Exit: every assessment sentence can point to an atomic record and source span.

### Phase 2 — Named-system deep dives (weeks 2–4)

- Vercel/eve, Fabro, Tessl, Mastra, SwarmForge, Agentic Coding Flywheel, HULA, Claude Tag/Agent Teams, Stakpak;
- add Factory/Cognition/GitHub/Cursor/OpenHands and internal enterprise factory exemplars;
- run stable-tag code audits where source exists;
- record “unknown” and disconfirmation queries instead of filling gaps by inference.

Exit: capability, human, Day-2, slop, adoption, and evidence profiles for the first 20 systems.

### Phase 3 — Day-2 and modernization lanes (weeks 3–5)

- assess HolmesGPT, OpenSRE, Resolve, Rootly, incident.io, Datadog, Kubernaut, Akmatori and related systems; retain AWS agent products only as explicitly labeled vendor-claim/negative-control cohorts until they earn qualification;
- assess AWS Transform, IBM watsonx, OpenRewrite/Moderne, Reversa, and other modernization systems;
- create incident-replay and semantic-equivalence fixtures.

Exit: explicit boundary between construction factories, factory operations, and application operations.

### Phase 4 — Benchmarks and pilots (weeks 4–8)

- reproduce public results where possible;
- run the same model across multiple factory scaffolds;
- run FactoryBench brownfield, slop, secure delivery, incident, and learning tracks;
- design an internal stepped rollout with telemetry and blinded reviews.

Exit: no benchmark claim lacks version, scaffold, artifacts, cost, and human-intervention data.

### Phase 5 — Publication and recurring job (weeks 6–9)

- generate system cards, comparison tables, wiki synthesis, and decision tools;
- add validation and reflect-audit fixtures for recurring overclaim patterns;
- install bounded scheduled discovery with dry-run, status, stop, and self-stop controls;
- maintain corrections as append-only inputs to future passes.

Exit: three no-material-finding passes trigger maintenance cadence, and the full corpus rebuilds deterministically.

## First analytical hypotheses

These are questions to test, not conclusions:

1. The best current factories optimize human review throughput, not full autonomy.
2. Specialized agents with typed, inspectable handoffs outperform one omnibus agent operationally even when they do not improve base-model capability.
3. Deterministic verification and realistic environments matter more than adding reviewer agents.
4. Most “self-improving” systems improve prompts and evals, not the application's architecture.
5. Day-2 operations remains a separate market from issue-to-PR factories, with weak handoff between the two.
6. Legacy modernization succeeds when behavior recovery, deterministic transforms, and domain approval surround generative work.
7. Public benchmark rankings currently measure coding harnesses more often than factories.
8. The binding constraint shifts from code generation to specification quality, review attention, environment fidelity, and operational authority.

## Immediate next actions

1. Implement the append-only schema in `ontology.yaml` and a validator.
2. Ingest the named-system sources from `research-queue.md`.
3. Clone open systems to temporary directories, check out latest stable tags, and produce source-backed capability cards.
4. Build a public-PR collector for Vercel AI SDK Factory and other bot identities.
5. Reproduce Fabro's SWE-bench Lite scoreboard and add a SlopCodeBench adapter.
6. Create the first FactoryBench legacy fixture from a deliberately tangled but observable service.
7. Interview or obtain artifacts for Dana Marble's system; keep its status provisional until then.
8. Decide whether Iron Reins should live inside `stateofai` as a new technical pillar or in a separate repository with a generated State of AI publication feed.
