# Iron Reins: enterprise software factories and autonomous technology organizations

Prepared: 2026-08-24  
Scope: enterprise systems that repeatedly turn intent, incidents or maintenance obligations into verified software and operational changes  
Evidence policy: first-party source and repository evidence preferred; vendor adoption counts labeled; unavailable/private implementation recorded as a gap

Commercial evidence is edition-specific. A public repository does not make a hosted control plane open source, and a zero-dollar license does not imply a zero-cost factory. The dated edition, gate, deployment and cost ledger is in [`commercial-boundaries.md`](commercial-boundaries.md).

## Executive conclusion

There is not yet one publicly evidenced factory that can safely take “move this unmanaged EC2 MLflow service to a production multi-tenant Kubernetes platform with our CI/CD templates, Okta, notebook authorization, fully tested sandbox/staging/production, and ongoing operations” and own the result unattended until completion. Iron Reins therefore begins with a narrower enterprise paved road: versioned Copier templates for EKS, fixed Terraform/Helm/GitOps building blocks, explicit policy and real-environment tests.

The current phase is research only. No candidate is being installed or run. Pass 002 adds IT4IT as the whole-lifecycle information-model backbone and an OSS assembly hypothesis centered on explicit authorities: an ITSM/work record, one catalog/platform API, Git and deterministic reconcilers, Okta identity, telemetry, draft-first agents and human authority. OpenChoreo, Kubernaut, Deputies, Amika, GLPI, Tessio and the KubeRocketCI testbed are research candidates—not recommendations.

Every candidate now receives an operator-signal smoke test before any future qualification budget. X, Reddit, GitHub issues and support forums can reveal low operator value, setup failures and scale limits much earlier than vendor material. The lane weights concrete environment/task/failure detail and corroboration; vague sentiment and unanswered “has anyone used this?” threads remain leads or adoption gaps. This correction keeps AWS agents in the negative-control cohort.

GitHub, Hugging Face and Kaggle are now separate evidence lanes. GitHub is authoritative for public implementation and issues; Hugging Face is useful for benchmark/model/dataset discovery; Kaggle is primarily a source of triage corpora and synthetic or small fixtures. New high-value leads include IncidentFox's explicit open-core security boundary, Aurora's multi-agent incident workflow, Keep's alert-management substrate, Evidra-Lock's deterministic fail-closed action gate, BeyondSWE's 178 dependency-migration tasks, AIOpsLab and AgentRx. Synthetic log datasets and model tags receive no causal or production credit.

This narrower sphere does not solve root-cause analysis. Diagnosis, remediation selection and recovery are scored separately. A plausible narrative receives no causal credit unless it predicts a discriminating observation; a successful restart or rollback is mitigation unless the underlying cause is established; and no repair receives safety credit unless desired state, rollback, collateral effects and recurrence are verified.

The market has converged on a stack of separable capabilities:

1. a durable control plane that decomposes work, persists state, enforces budgets and routes approvals;
2. coding agents operating inside reproducible isolated environments;
3. deterministic templates and mutation engines for infrastructure, identity, builds, package resolution and deployment;
4. independent acceptance systems—tests, policy, browser/API evidence, security, architecture and model judges;
5. real target-system qualification environments, not only generic sandboxes;
6. observability/RCA/remediation agents for Day 2;
7. humans retaining authority over ambiguous product intent, privileged identity, data semantics, production cutover and irreversible actions.

The practical recommendation is to build a compositional factory and evaluate it against brownfield FactoryBench scenarios. “Agent ran for a long time” is not the objective. The objective is a convergent, auditable system whose result survives independent verification, deployment, incidents, upgrades and organizational handoff.

The same standard now applies to time. Every material mutable fact in a greenfield paved-road system preserves both business-valid time and recorded/system time, backed by a transactional semantic outbox, independent committed-change capture and a queryable bitemporal projection. For legacy systems, network monitoring contributes observation history but never becomes invented transaction or valid-time truth; snapshots, audits, CDC and semantic events progressively raise a separately recorded capture-fidelity grade. See [bitemporal enterprise information and CDC](ecosystems/bitemporal-information-and-cdc.md) and [`fb-temporal-truth-001`](benchmarks/bitemporal-enterprise-information.md).

Temporal truth also extends through consumption. Iron Reins requires reconstructable `decision/action → data-use event → invocation → derived field → transformation → upstream field versions → source fact versions` lineage. “As of” must identify its clock: valid, known/recorded, commit, observation, computed, fitness-evaluated, used and decision-effective times are distinct. A recent model response or cache entry does not prove that its underlying data was fresh or fit for a particular decision, and a later correction must enumerate the exact calls, decisions and actions that consumed the affected versions.

Commercial technology is now modeled as its own operating organ. Procurement/contracts, invoices/cards, IdP/SCIM, native SaaS/COTS administration, endpoint inventory, cloud/FOCUS billing and usage observations remain independent evidence sources reconciled into bitemporal vendor/product/SKU, agreement, entitlement, assignment, use, service/value, renewal and action records. OpenCost, OptScale, Infracost, Cloud Custodian, GLPI/Snipe-IT, Fleet/OCS, midPoint/Syncope and OSS-license scanners cover useful slices; no mature OSS system or agent found closes the heterogeneous contract-to-reclaim or observation-to-realized-savings loop. Vendor collectors receive acquisition credit only until native reconciliation and outcome proof. See [FinOps, commercial assets and research promotion](ecosystems/finops-commercial-assets-and-research-promotion.md) and [`fb-commercial-ops-001`](benchmarks/commercial-operations-and-research-promotion.md).

The same pass rejects a broad claim that TUIs are replacing Jupyter. The supported interpretation is hybridization: Jupyter remains important for exploration, while Jupytext and marimo improve text/source behavior, packages and locked environments carry reproducible logic, Dagster owns production data execution, and CLI/TUI layers remain replaceable operator adapters. A terminal interface is not a production architecture.

OpenAI's Kepler adds a strong data-agent reference pattern, not a complete data department. Its six context layers combine catalog/query-history/lineage metadata, human annotations, Codex-derived code semantics, permission-aware institutional knowledge, scoped memory and live warehouse/Spark/Airflow/metadata calls. Primary evidence supports production use, roughly 70,000 cataloged datasets and more than 600 PB/day processed by the underlying data platform. It does not support transferring 3.5k platform users into active Kepler users or generalizing one 22m41s-to-1m22s repeat query into a fleet latency claim.

OpenMetadata is not a new Kepler-era component: its repository and first public non-prerelease release date to August 2021, with lineage, pipeline ingestion and profiling listed by October 2021 and metadata version/change events by November 2021. This five-year public lineage strengthens the maintenance prior but does not establish five years of enterprise production. Iron Reins separately tests whether its schema/standard layer can be decoupled and round-tripped from the implementation without semantic or control loss.

The current OpenMetadata `1.13.4` source also creates a P0 qualification gate. Keyword MCP search passes a user subject into RBAC-aware search, but semantic MCP search directly queries the vector service without using its authorizer/security inputs; search access control is disabled by default. Open issue `#30023` reports this semantic-domain leak class at runtime on `1.12.5`, but Iron Reins has not reproduced it on `1.13.4`. Every retrieval/index/cache path, bot/user/delegated identity, revocation and memory boundary remains a sealed test in [`fb-data-agent-context-001`](benchmarks/data-agent-context-and-permission-fidelity.md). OpenMetadata licensing is also split: Apache-2.0 core, but source-available Collate Community License for official MCP and ingestion modules. See the full [Kepler/OpenMetadata dossier](ecosystems/kepler-openmetadata-data-agent.md).

OpenLineage, OpenMetadata and OpenBytes occupy different layers. OpenLineage `1.52.0` is an Apache-2.0 event specification, client and integration ecosystem whose run/job/dataset facets are producer assertions; it is not a durable event log, reconciled truth graph, source-policy engine or value-level provenance system. OpenBytes was an open-dataset schema/index project, moved to LF AI & Data archival status in May 2025 and now excluded as an active component. Iron Reins now models the full lineage-assertion supply chain and tests it with [`fb-lineage-context-001`](benchmarks/lineage-context-and-provenance-fidelity.md). See the [prior-art dossier](ecosystems/open-lineage-and-data-context-prior-art.md).

Data quality is now modeled as a separate operating discipline. A contract or expectation is a promise; a versioned observation records what actually ran against which population; an incident records impact and evidence; a repair records correction, replay/restatement and independent verification; and a waiver records scoped, expiring accepted risk. GX Core, Deequ, Soda Core, Elementary OSS, Pandera, `datacontract-cli` and Evidently cover useful but non-equivalent slices. Soda Core is ELv2 source-available rather than OSS, and paid observability/control-plane features must not be transferred to community components. [`fb-data-quality-001`](benchmarks/data-quality-operations-and-repair.md) tests false-green checks, late/duplicate/missing/corrected data, poisoned baselines, false alerts, incomplete lineage, bad repairs and downstream closure. See the [data-quality operating-system dossier](ecosystems/data-quality-operating-system.md).

AutoML is not dead; it decomposed. AutoGluon/FLAML/TPOT/H2O generate challengers, Optuna/Ray Tune/Katib run search, Pandera/GX/Deequ/TFDV validate inputs, and Evidently/NannyML observe production. None owns semantic correctness or safe Day-2 repair. H2O's announced migration of several deployment/runtime surfaces into H2O-3 Secure, the end of public GX Cloud, Soda Core's ELv2 license and SageMaker Model Monitor's closure to new customers are material boundary corrections. See the [AutoML survivor map](ecosystems/automl-and-ml-data-quality-survivors.md).

NVIDIA is the broadest public AI-factory substrate in the corpus, not a whole IT department. Its public estate spans kernels/compilers, RAPIDS data processing, Curator, Megatron/NeMo training and RL, Evaluator/Gym/garak, TensorRT/Triton/Dynamo serving, OpenShell/Relay/NemoClaw agents, and DSX/GPU/DCGM/NVSentinel fleet operations. The stack mixes true OSS, proprietary binary dependencies, artifact-specific model/data licenses, tested containers and paid AI Enterprise production branches. DSX Exchange is an event spine rather than transactional authority; NVSentinel is strong bounded GPU Day 2 rather than general causal RCA. See the [NVIDIA dossier](ecosystems/nvidia-software-and-model-factory.md) and [organization census](evidence/nvidia-github-organization-inventory.md).

The recent-turn integrity pass also corrected four omissions. Netlify Agent Runners is a material hosted web-factory comparator with isolated runners, previews and human deploy authority; Graphwise is a commercial semantic-context platform with open adapters, not a source-of-truth oracle; Sentra is a proprietary DSPM observation/remediation-proposal organ; and SQL Sentry is mature SQL operations telemetry whose own overhead and upgrade failures must be benchmarked. They are placed in the [recent enterprise platforms dossier](ecosystems/recent-enterprise-platforms-and-control-surfaces.md), while the [coverage ledger](passes/2026-08-24-recent-turns-coverage-ledger.md) records which recent research questions are fully or partially ingested.

For an enterprise already standardized on Microsoft, this conclusion is even stronger. HR, hybrid AD DS, Entra, licenses, Exchange, Teams, SharePoint/OneDrive, Intune/Windows, Purview, Power Platform and Azure form a distributed administrative system, not one API. A completed workflow or successful Graph/PowerShell/DSC action is not a completed employee or business-user outcome. The new [Microsoft enterprise operations organ](ecosystems/microsoft-enterprise-operations-organ.md) requires field-level authority, expiring privilege, cross-workload compensation, source-native reconciliation, user-journey probes, entitlement-aware audit and a vendor-change radar; [`fb-microsoft-estate-001`](benchmarks/microsoft-estate-autonomy.md) tests the assembly under throttling, sync conflict, stale export, legal hold, offline endpoint, license and API/module failures.

Commercially, the market is also not one category. Warp exposes substantial public worker/configuration surfaces but keeps the central service proprietary; Vercel Open Agents is an MIT reference whose required managed primitives are separately metered; Fabro is MIT and self-hostable with adopter-owned operations; Mastra combines an Apache framework and pre-1.0 Factory package with a $0/$250/custom managed platform; Tessl, Ona, Factory and Cursor reserve important governance or control-plane functions for paid tiers; internal systems such as Minions, Inspect and Honk are not products with public prices. These boundaries materially change reproducibility, security responsibility and TCO.

## Market map

| Category | Strong examples | What is actually proven |
|---|---|---|
| internal PR factories | Stripe Minions, Spotify Honk/Fleetshift, Ramp Inspect, Coinbase Mux | organization-scale parallel change generation with human merge; strongest for bounded, testable work and fleet migrations |
| open/reference cloud factory platforms | Warp Factories/Oz, Vercel Open Agents/eve, Fabro, Mastra Factory | durable orchestration, isolated execution, workflow-as-code and integrations; maturity and open/closed boundaries vary |
| local multi-agent factories | SwarmForge; NTM/Agent Flywheel; Claude Agent Teams | role decomposition, local orchestration, work ownership and repository checks; operator skill and project adapters remain decisive |
| agent/workflow substrates | Block Goose, Shopify Roast, OpenCode | flexible loops and tools used to build factories; not sufficient factory governance alone |
| Day-2 operations systems | Ramp's Sheets loop, Stakpak, HolmesGPT, OpenSRE and other reproducible OSS operators | telemetry-led investigation and bounded remediation; Ramp is the clearest public production-signal-to-fix-proposal example among coding factories; managed vendor agents require operator proof before inclusion |
| deterministic modernization | OpenRewrite/Moderne, AWS Transform, IBM watsonx Code Assistant for Z | repeatable mass change and specialized legacy transformation; general semantic preservation and organizational redesign remain human-led |
| enterprise administration | Microsoft-native agents over M365/Entra/Intune/Azure plus Microsoft365DSC, Maester, Graph/PowerShell, Update Manager and Autopatch | broad governed coverage when composed; no safe single agent spans cloud, tenant, endpoints, on-premises AD and all rollback domains |

## How far the leading systems have gotten

- **Stripe Minions:** strongest published internal one-shot PR throughput—more than 1,000 wholly agent-produced, human-reviewed merged PRs per week—backed by isolated devboxes, curated tools, conditional rules, selective CI and bounded repair. No public Day-2 loop, implementation, model matrix or defect study.
- **Spotify Honk:** strongest published fleet-migration and verifier story—1,500+ merged Honk PRs, repository-sensitive checks, stop hooks and an independent judge. Failed migrations show that missing context and tests remain hard boundaries.
- **Ramp Inspect:** broad internal background-agent adoption and the best documented Day-2 bridge. Generated monitors can launch sandbox reproduction and a proposed fix, but humans author trusted instrumentation and merge; noisy thresholds and duplicate alerts remain real failure modes.
- **Vercel AI SDK Factory:** unusually transparent production loop for one open-source product, with reported 25–35% of weekly merged PRs and over 75% of a month's issue closures. Its success depends on custom bug probes, live E2E evidence, independent review and human merge.
- **Warp Factories:** broadest productized control-plane ambition: workflow-as-code, cloud workers, environments, skills, checkpoints and evolving-factory feedback. Public components are substantial, but the central control plane remains proprietary and application Day-2 closure is not yet demonstrated publicly.
- **Open Agents:** credible open Vercel-native cloud-agent reference, not an enterprise factory out of the box. It supplies durable execution, sandbox lifecycle and optional PR creation but leaves acceptance policy, enterprise authorization and operational closure to adopters.
- **Fabro:** strong open/self-hostable deterministic workflow candidate with sandboxing, checkpoints and evaluation orientation. It is an orchestrator, not a turnkey identity/deployment/SRE system.
- **SwarmForge and NTM/Agent Flywheel:** useful local factory ecosystems whose language breadth comes through project adapters and attached tools. Their strongest value is operator-visible coordination and verification; their factory-level external benchmark evidence is limited.

Full comparisons: [`README.md`](README.md), [`repository-index.md`](repository-index.md), [`ecosystems/internal-enterprise-factories.md`](ecosystems/internal-enterprise-factories.md), [`ecosystems/open-agents.md`](ecosystems/open-agents.md), and [`ecosystems/warp-factories.md`](ecosystems/warp-factories.md).

## Day 2, slop and long-horizon evidence

Most “software factories” end at a branch or pull request. Sandbox lifecycle is Day 2 for the agent runtime, not Day 2 for the produced application. A qualifying Day-2 loop starts with live telemetry or an operational obligation, reconstructs evidence, proposes a bounded action, obtains the required authority, verifies the live outcome, preserves rollback, and turns the incident into a regression fixture.

Ramp supplies the strongest public example in the internal-factory cohort. Stakpak, HolmesGPT and OpenSRE provide independently inspectable adjacent operational pieces. AWS DevOps Agent has been demoted to a vendor-claim/negative-control cohort after direct negative operator feedback; no managed agent is recommended without sealed-trial and customer-outcome evidence. None establishes a universally safe autonomous remediation system.

No complete named factory in this review has a verified SlopCodeBench result. Underlying-model or agent scores must not be inherited by the factory. SlopCodeBench is valuable because it measures structural erosion and verbosity over iterative changes, but a factory benchmark must also measure acceptance-policy integrity, test tampering, reviewer load, deployment survival, incident behavior, maintainability and rollback. The external benchmark reports that no evaluated agent completed a full problem end to end and that many trajectories accumulated structural or verbosity degradation. See the [paper](https://arxiv.org/abs/2603.24755) and [leaderboard](https://www.scbench.ai/).

The most credible anti-slop pattern is layered:

- constrain scope and express intent as executable contracts;
- give builders hermetic environments and only required tools;
- run deterministic repository checks before model review;
- use an independent verifier with authority to reject or request evidence;
- test behavior, architecture, security, performance and deployment—not just syntax;
- preserve failed trajectories and production incidents as future fixtures;
- measure reverts, escaped defects, complexity/coupling drift and human review time after merge.

## Cloudflare and qualification capacity

[Cloudflare OS](https://github.com/cloudflare/cloudflare-os) is a real first-party open-source project launched in August 2026. It is an isolate-based workspace for generated Worker-compatible JavaScript/TypeScript “Gadgets,” with Durable Object state and Gatekeeper-controlled external actions. It is useful for the qualification portal, lightweight evaluators and human approval surfaces, but it is not a full Linux software-factory runtime.

The general coding tier is Cloudflare Sandbox SDK/Containers, with Cloudflare Agents or another durable scheduler. Cloudflare Environments for Claude Managed Agents is a Claude-specific integration rather than a model-neutral factory brain.

Cloudflare is a strong high-fan-out evaluation fabric. Published platform limits include 1,500 concurrent vCPU, 6 TiB memory and 30 TB disk, with up to roughly 15,000 small instances; an individual container is much smaller unless negotiated. This is enough for many parallel repository builds, tests, policy checks, browser/API fixtures, traces and durable evaluators. Modal is a useful alternative, including thousands of sandboxes and full-VM beta when a real kernel is needed.

Neither generic fabric proves target-cloud semantics. The MLflow scenario still needs disposable AWS accounts/roles and real EKS for IAM, VPC, ALB, KMS, RDS, S3, EBS/CSI, DNS and node-disruption behavior. Microsoft/AD changes need disposable Windows VMs and domain controllers. A sensible first qualification cell is 50–100 concurrent workers; model quotas, test fixtures, evidence review and target-cloud limits are likely to bind before raw sandbox count.

Details: [`infrastructure/factory-qualification-lab.md`](infrastructure/factory-qualification-lab.md).

## MLflow migration as a comparative stress test

MLflow is used here as a benchmark mission, not as a proposed implementation project. The comparison asks how much of the mission each ecosystem can complete, what it must import, which evidence it can generate, and where a human or specialist system must take over. The full ecosystem matrix is in [`factorybench/mlflow-ec2-to-multitenant-kubernetes.md`](factorybench/mlflow-ec2-to-multitenant-kubernetes.md).

The capabilities a complete solution would have to compose are:

| Plane | Preferred pattern |
|---|---|
| work control | Warp/Fabro/NTM-style durable graph with explicit owners, retries, budgets and approvals |
| implementation | interchangeable Claude/Codex/Warp/other agents selected by controlled task tests |
| fast evaluation | Cloudflare or Modal sandboxes for parallel builds, policy, synthetic services and evidence |
| target truth | disposable EKS and short-lived AWS roles/accounts; real Okta test application/tenant path |
| authoritative mutation | internal Terraform/Helm/GitOps templates, MLflow tools, Okta Terraform/API, native resolvers |
| verification | deterministic tests plus independent review; Stakpak/Holmes/OpenSRE-class operational checks, with managed vendor agents admitted only after qualification |
| production authority | platform, identity, security, database, ML service and incident owners |

The largest migration risks are metadata/artifact preservation, the tenancy boundary, and notebook identity—not manifest generation. MLflow workspaces are not hard isolation. Existing artifact URIs do not automatically follow a new default store. The FileStore migration command is not a turnkey SQLite-to-PostgreSQL plus artifact migration. Database upgrades can be nontransactional. Okta support commonly depends on a community OIDC plugin or proxy, and notebook token refresh/revocation needs an explicit broker or Kubernetes service-account design.

The pass/fail contract requires reconciliation of every promised experiment/run/model and artifact, cross-tenant denial, Jupyter authentication, backup restore, disruption, staged upgrade, rollback and a live Day-2 exercise before completion. No named ecosystem currently passes all dimensions on public evidence.

## Whole-SDLC benchmark strategy

No single benchmark spans product discovery, sustained development, review, deployment, production incidents and months of evolution. Iron Reins therefore uses a vector of component benchmarks plus 30-day shadow and 90-day canary trials. The stack includes METR Time Horizon/HCAST, SWE-bench variants, SWE-EVO, RoadmapBench, SWE-Chain, EvoClaw, SlopCodeBench, DevOps-Gym, c-CRAB, SecureAgentBench, AIOpsLab, ITBench, SREGym, OpenSRE and DORA outcome telemetry.

The complete protocol, metrics and contamination controls are in [`benchmarks/whole-sdlc-benchmark-stack.md`](benchmarks/whole-sdlc-benchmark-stack.md).

## Humans: where they remain necessary

Humans should supply or approve:

- product intent and behavior where no executable oracle exists;
- tenant isolation, data classification, RTO/RPO and cost boundaries;
- privileged identity consent, role/group assignment and break-glass policy;
- production schema migration, destructive data operations, DNS/traffic cutover and rollback decisions;
- exceptions where security, compliance or architectural baselines conflict;
- acceptance when tests are incomplete or the change alters business semantics;
- retirement of the old system after the evidence and rollback window are complete.

The factory can earn narrower authority over repeated low-risk classes. Authority should be promoted by observed success, not by a global autonomy setting.

## Research machine

Iron Reins deliberately mirrors the stateofai pattern: versioned ontology, append-only evidence and corrections, source boundaries, disconfirmation, periodic replay, coverage queries and bounded scheduled runs. It separates source artifacts, systems, releases, repository components, language capability, model compatibility, lifecycle evidence, human gates, deployments, benchmarks, failures, corrections, queries and decisions.

Core artifacts:

- [`ontology.yaml`](ontology.yaml) — canonical concepts and validation rules;
- [`dossier-template.md`](dossier-template.md) — required per-ecosystem evidence fields;
- [`research-queue.md`](research-queue.md) — unresolved systems and tests;
- [`README.md`](README.md) — ontology, scoring, recurring process and initial market map;
- [`repository-index.md`](repository-index.md) — canonical/adjacent/private repository boundaries.
- [`commercial-boundaries.md`](commercial-boundaries.md) — stable-tag status, exact commercial gates, control-plane ownership, limits, dated price observations and the `S/M/C/I/P/X/H` cost ledger.

## Interactive Greenfield reference architecture

The interactive Atlas now carries one deliberately opinionated composition rather than only a market map. For a new US-first company below 300 employees, the reference starts with AWS Organizations and EKS, OpenTofu for cloud resource plans, Argo CD for Kubernetes desired state, GitLab Premium with independent ephemeral runners, RDS PostgreSQL as the default transactional record store, S3 for objects and immutable evidence, Dagster plus dbt Core for data assets, OpenMetadata plus OpenLineage for context and runtime lineage, OpenTelemetry plus the Grafana stack for signals, and cross-account AWS Backup/Velero/Kopia recovery.

The workplace baseline is Microsoft 365 Business Premium because the published US annual list price is $22 per user per month and the bundle includes Entra ID P1, Intune and Defender for Business. Microsoft explicitly documents full Intune capabilities for macOS as well as iOS/iPadOS, Android and Windows; Jamf is therefore a measured-gap addition, not a Greenfield default. [Microsoft 365 pricing](https://www.microsoft.com/en-us/microsoft-365/business/microsoft-365-plans-and-pricing), [Business Premium security and Intune FAQ](https://learn.microsoft.com/en-us/microsoft-365/admin/security-and-compliance/m365bp-security-faq?view=o365-worldwide)

Cloudflare Zero Trust supplies the initial workforce edge: its public plan is free through 50 users and the pay-as-you-go plan is $7 per user per month, but log retention, support and full SASE/DLP depth differ by plan. GitLab Premium is modeled at its published $29 per user per month annual price for engineering seats. 1Password Business is modeled at $8.99 per user per month annually, with its ten-person Starter Pack represented separately. [Cloudflare Zero Trust pricing](https://www.cloudflare.com/plans/zero-trust-services/), [GitLab pricing](https://about.gitlab.com/pricing/), [1Password Business pricing](https://1password.com/pricing/business)

The data control-plane decision keeps Dagster OSS available while making the managed boundary visible. Dagster+ Starter is publicly priced at $100 per month plus $0.035 per credit and $0.010 per serverless compute minute; Pro gates SSO/SAML/SCIM, audit logs, granular roles and enterprise support behind custom pricing. OpenMetadata remains Apache-2.0 and supplies catalog, lineage, quality and agent context, while Collate's managed and enterprise boundaries are assessed separately. [Dagster pricing](https://dagster.io/pricing), [Dagster repository](https://github.com/dagster-io/dagster), [OpenMetadata repository](https://github.com/open-metadata/OpenMetadata), [Collate pricing](https://getcollate.com/pricing)

PostgreSQL is the default, not the universal substrate. Relational control objects, bitemporal facts, contracts and metadata links belong there. Large immutable artifacts and evidence belong in S3; high-volume logs, metrics and traces remain in purpose-built telemetry stores; search indexes are derived and rebuildable. EKS itself costs $0.10 per cluster hour under standard Kubernetes support before nodes, volumes, public IPv4 addresses and traffic. RDS, nodes, telemetry and model usage are therefore estimates in the interactive cost model, not false precision. [Amazon EKS pricing](https://aws.amazon.com/eks/pricing/), [Amazon RDS for PostgreSQL pricing](https://aws.amazon.com/rds/postgresql/pricing/)

The composition deliberately defers Crossplane: OpenTofu owns AWS resource state and Argo CD owns Kubernetes desired state. Crossplane is agent-legible and powerful, but adding its providers, CRDs, functions, finalizers and continuous reconciliation before a concrete self-service need would increase the controller and recovery estate. It is a future trial, not a rejection. [Crossplane repository](https://github.com/crossplane/crossplane), [OpenTofu repository](https://github.com/opentofu/opentofu), [Argo CD repository](https://github.com/argoproj/argo-cd)

No cost total includes quote-only MDR, advanced IGA/DLP, regulated legal advice, an external audit opinion, physical endpoint logistics or accountable incident command. Those are visible uncovered functions. The reference stack can automate evidence, proposals, reversible actions and coordination; it cannot honestly create a zero-accountability company.

## Limitations and next evidence

This is a source-backed baseline, not a completed census. Several important systems are private, new or described only in vendor material. Model behavior and product surfaces change rapidly. GitHub activity and runtime presence do not prove language competence. Public PR counts lack common denominators. Day-2 claims require production telemetry and outcome studies.

The next phase should pin every public repository to a commit, populate append-only evidence records, interview private-factory operators, and run the same FactoryBench corpus across candidate control planes. The MLflow scenario should be first because it exposes code, cloud, identity, data, security and operations failure modes in one auditable task.
