> ## Content Index
> Fetch the complete content index at: https://www.shiftharness.tech/llms.txt
> Use this file to discover other available public pages before exploring further.

# The Harness Has a Cost: Which AI Delivery Controls Earn Their Place
- URL: https://www.shiftharness.tech/ai-generated-code-delivery-controls/
- Published: 2026-09-11T19:59:38.000Z
- Updated: 2026-09-11T19:59:39.000Z
- Description: Every serious piece this year told you to add one more AI delivery control. None told you which one to remove. Here is a subtraction test for the controls you already funded.
- Author: Sergii
- Tags: AI Enablement, Delivery Teams, AI in Software Delivery, AI Adoption

You already funded the harness: the prompt library, the repository context file, the architecture board, the review checklist. Almost every serious piece published this year told you to add one more. None told you which one to remove. Most also treat a green build as verification, which is the assumption that makes the whole stack look cheaper than it is.

Here is the subtraction test. Where qualified verification attention has become the binding constraint, a discretionary delivery control earns its place only if it makes trustworthy change cheaper to verify. It must reduce attention at the same risk coverage, widen coverage at the same attention cost, or increase verified throughput without shifting the cost into escaped defects, delivery latency, or lost system understanding.

Controls required by law, segregation of duties, or catastrophic-risk policy answer to another authority. The test cannot authorize their removal. It can still expose an unnecessarily expensive implementation.

The test is deliberately narrow. It should disqualify things you have already built.

Addition has structural advantages over subtraction. Vendors can package a new control, platform teams can own it, auditors can see it, and leaders can announce it. Nobody receives equivalent credit for removing a control whose prevented failures are invisible. Loss aversion and blame asymmetry therefore make the harness much easier to grow than to prune.

Adding controls to govern **ai generated code** is easy. Building controls that increase delivery capacity is a different job, and the two get confused.

## The test has to survive contact with arithmetic

A test you cannot compute is a slogan. A test you compute badly is worse, because it produces a number people defend.

**Calculate within risk tiers, and never aggregate first.** For each tier r:

> Verified attention efficiency (r) = (eligible changes in tier r whose required evidence was accepted for the revision that shipped, counted in a fixed change unit) / (qualified verification attention + amortized control-maintenance attention)

Compare routine to routine, elevated to elevated, boundary to boundary. A single weighted ratio across the portfolio is unsafe: a team improves it by shipping riskier work, by splitting changes differently, or by reclassifying them into higher-weighted tiers. Conventions like 1 for routine, 3 for elevated, 10 for material-boundary are fine for planning assurance effort, but they cannot declare a control effective. For an executive aggregate, weight tier results against a fixed baseline portfolio mix, never the current period's. Fix the protocol before comparing periods: which changes are eligible, how rejected and reworked changes count against attention, how shared maintenance is allocated, how long the observation window runs, and whether the numerator counts release assurance or completed operational validation, reported separately rather than mixed.

> **Evidence arrives at two checkpoints.** Release assurance is what you have at the release decision: specification-derived acceptance tests, domain invariants, security properties, contract tests, fitness functions, rollback evidence. Operational validation is collected in a defined post-release window: canary behavior, error-budget impact, rollback thresholds, business-invariant checks. A change can be sufficiently assured for release without yet being operationally validated, and a formula that counts production signals before the change has had real exposure is counting evidence that doesn't exist yet.

> **The denominator is broader than review hours.** Where verification capacity is binding, qualified human attention is the primary denominator. Control latency, direct tooling cost, and downstream quality remain guardrails, and a control that improves the ratio while degrading any of them has moved the cost rather than removed it. Separate setup cost from operating cost too: a fitness function can be expensive for eight weeks and nearly free for two years afterwards, so judging it inside the installation window rejects precisely the controls that amortize best.

> **Before-and-after is a screening signal, not a control-effect estimate.** Model capability, team composition, risk mix, release volume, batch size, and agent familiarity can all move in the same window. To estimate delivery impact, prefer staggered rollout or matched comparison. To test whether the control catches what it claims, use incident replay, seeded violations, or shadow mode. Treat before-and-after trend as the weakest evidence.

> **Controls interact, so evaluate marginal contribution.** A specification becomes valuable because it drives an oracle; a risk classifier because it routes to an owner; human review because automated evidence narrowed the question. A review gate can look unproductive only because another control absorbed its detections. The test asks what risk coverage, attention cost, and latency change when a control is added, narrowed, shadowed, or removed.

That gives the criterion its final form.

> Where qualified verification attention is the binding constraint, a discretionary control earns its place only through measurable marginal contribution: it must reduce the attention required to achieve the same risk coverage, increase risk coverage at the same attention cost, or improve verified throughput without worsening escapes, latency, or system understanding.

> Mandatory controls answer to an external authority and cannot be removed by this test, but their implementation cost and efficiency should still be measured. Controls should be evaluated within comparable risk tiers, against a fixed change mix, and as part of the control portfolio around them.

That last clause matters more than it sounds. Otherwise "mandatory" becomes a shield around an inefficient implementation nobody is allowed to improve.

![Three tier worksheets headed VERIFIED ATTENTION EFFICIENCY, labelled ROUTINE, ELEVATED and MATERIAL BOUNDARY, each worked to a different result; a PORTFOLIO AGGREGATE sheet lies blank.](https://storage.ghost.io/c/73/3e/733efc05-c397-4cf1-b19b-527c1b07dee7/content/images/2026/09/image-2-2.png)

## Two terms carry the weight

Compiling, a passing suite, and an approving reviewer are three facts about a change. None of them, on its own, is verification.

A change is verified when there's credible evidence that it produces the intended behavior, respects the constraints that apply to it, leaves critical system properties intact, can be operated safely, and stays reversible where reversibility matters. That is verification in the broad sense, spanning what IEEE 1012 separates into verification against specified requirements and validation against intended use. Verification is the accumulated evidence behind a decision to release and keep a change. Treating it as a single gate is what lets a team believe the gate is the verification. The [cost of verification](https://www.shiftharness.tech/cost-of-verification-ai/) is the argument underneath this one. This piece assumes it and asks the narrower question of what to do about the controls already funded.

The scarce input is qualified attention, not engineering hours. Somebody with enough context to spot a bad assumption, follow a structural implication, and take responsibility for a high-impact change. A senior architect's thirty minutes is not interchangeable with thirty minutes of generic review capacity. Headcount is a poor substitute. Generating more code doesn't generate more of that.

Moderne argues the adjacent point well: match scrutiny to the risk of the change. But that allocates scrutiny across changes already flowing through a control. This test asks the prior question, because a team can allocate scrutiny beautifully across a gate that should never have been built.

## One enterprise just ran the mandate, and the result is uncomfortably on the nose

In July 2026, He, Agarwal, Denisov-Blanch, Azaletskiy, Koyejo and Vasilescu published a longitudinal case study of a documented enterprise "2x" mandate, titled *AI Writes Faster Than Humans Can Review*. The data covers 802 developers and 196,212 pull requests between January 2024 and April 2026\. Among the 564 developers observed for at least three active months, per-capita throughput reached 2.09 times the pre-mandate baseline, rising from 21.2 authored pull requests per active developer in the January to April 2025 baseline window to 44.3 by April 2026\. That is among the largest gains reported from a field deployment of AI coding tools.

The second half of the abstract is worth reading twice. Adoption restructured code review around automation: per-reviewer load roughly doubled, automated review overtook human review, and merge and revert rates held steady. That the [review cost moves rather than disappears](https://www.shiftharness.tech/ai-code-review-cost-shift/) is well documented by now. What this case study adds is a measured shape for the move.

Read the caveats with it, because the authors state them plainly. This is one mid-sized, AI-forward company, and adoption was not randomly assigned, so the authors read their staggered difference-in-differences design as strongly implicating an adoption-and-use channel rather than exact causal attribution, with the mandate acting as a catalyst rather than a direct driver. They are equally direct about the quality signals: merge and revert rates are, in their words, coarse, short-horizon proxies that miss defects, incidents, and maintainability. The extension is mine rather than theirs, but it matters: the export failure described later would not have moved either number for two quarters.

DORA supplies the broad organizational view. Its 2024 research associated a 25% increase in AI adoption with an estimated 1.5% reduction in delivery throughput and a 7.2% reduction in delivery stability. The 2025 report, based on roughly 5,000 technology professionals, found adoption had become positively associated with throughput and product performance, while the negative association with stability persisted. In that same population, 90% reported using AI at work and around 30% still reported little or no trust in AI-generated code.

Throughput moved. Stability did not follow it.

METR warns against trusting perception. Its July 2025 randomized study of sixteen experienced open-source developers working real issues in repositories they knew well found they took 19% longer with AI available, while afterwards estimating AI had made them roughly 20% faster. That is a snapshot of one population on one class of task, not a standing finding. The February 2026 follow-up is less quotable and more instructive: 30% to 50% of participating developers said they chose not to submit some tasks because they did not want to do them without AI, the time estimates moved toward a speedup (minus 18% for returning developers, minus 4% for new recruits), and every confidence interval crossed zero. Selection effects were large enough that METR treated the estimate as unreliable and began redesigning the study.

Self-report still earns its place; it just can't be the sole estimate of objective productivity. Surveys explain experience and mechanism. System data tests whether perceived speed translated into observable outcomes.

The **ai code review bottleneck** doesn't yield sustainably to adding generic reviewers, because the scarce capacity is repository-specific judgment, and distributing a diff across more people doesn't manufacture that context. Margaret-Anne Storey's triple-debt model explains why: technical debt lives in the code, cognitive debt lives in people as shared understanding erodes, and intent debt lives in externalized knowledge when goals, constraints, and rationale are never captured well enough for the next human or agent to use. Practice is heavily optimized for the first, and decades of linters, type systems, and static analysis all point at code.

AI-assisted delivery creates conditions in which the other two may grow faster: implementation expands before shared understanding and externalized rationale are replenished. Work arrives syntactically clean, well tested against its own interpretation of the requirement, and functional in production, while fewer people can say why it's correct. The code passes while understanding thins out. That is one plausible mechanism for the stability lag, not an established cause. It also explains why documentation disconnected from executable controls may fail to improve delivery outcomes: the artifact exists, but it does not constrain behavior at the point of change.

![Three stacked inspection plates: only the bottom one is etched with static-analysis output naming files, lines and rules, above a brass STATIC ANALYSIS label; the top two are blank.](https://storage.ghost.io/c/73/3e/733efc05-c397-4cf1-b19b-527c1b07dee7/content/images/2026/09/image-3-2.png)

## An oracle is only worth what its provenance is worth

This is the load-bearing component, and where much **ai for code review** enthusiasm quietly goes wrong.

When the same model interprets the requirement, writes the implementation, and generates the tests, the failure modes are correlated. A misunderstood requirement produces code implementing the wrong behavior, tests confirming that same wrong behavior, and an explanation justifying both with complete confidence. Shared-oracle coverage doesn't mean higher risk; it means the coverage number provides weak assurance, because the thing measuring correctness inherited its definition of correctness from the thing being measured. On a dashboard that looks like health: test count rises, coverage rises, escaped-defect rate doesn't move.

Temporal order is not independence. If one model interprets an ambiguous requirement, writes the specification, generates the test, and writes the implementation, the test existed first and all four artifacts inherit the same misunderstanding. Independence is a property of provenance, at three levels.

| Level                    | Meaning                                                                  | Example                                                                                                                                                 |
| ------------------------ | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Independent truth source | Correctness originates outside the implementation process                | Regulatory rule, customer contract, domain invariant, external API contract                                                                             |
| Independent derivation   | A different actor or process translates that truth into checks           | Domain owner writes authorization examples before implementation                                                                                        |
| Independent execution    | The check runs outside the implementation's own assumptions and fixtures | Provider-verified contract, replay of real traffic against separately owned expected results, property test whose invariants come from the truth source |

Pre-implementation tests are stronger evidence when their definition of correctness originates outside the implementation path. Temporal separation alone is insufficient. A customer-export implementation shouldn't define its own authorization truth; the property belongs outside it, and survives it. A second model or an adversarial prompt still helps, but those are diversified reviewers rather than independent oracles: different models share training data and similar blind spots, so they reduce correlation without eliminating it.

The design question isn't who wrote the test. It's where the test's definition of correctness came from.

## Six control capabilities most likely to earn their place

These describe capabilities, not products. Each is a prediction to test against your own harness. This article did not measure them.

> **Consequential specification.** A specification doesn't have to be executable. It has to be consequential. Compare "export customers to CSV" with a version specifying that only account administrators may initiate the export, that records stay within the account boundary, and that an audit entry is emitted. The second yields checkable assertions. Implemented as tests or policy checks and made a required gate, it can fail a build.

> **Codified workflow.** Prompt libraries depend on one person remembering which prompt applies, when, with what context. A codified process defines how work moves through repository instructions, CI policies, ownership rules, architecture checks, and incident-learning routines. An informal process scales through heroics; a codified one scales through repeatability, and once visible its performance becomes measurable.

> **Independently sourced oracle.** Covered above, and the one whose provenance you should be able to state out loud for any control you currently trust.

> **Evidence-based risk routing.** Uniform review wastes scarce engineering attention on changes that don't need it, but routing works only on evidence. A checkbox asking the author to declare a change low-risk provides weak assurance. Derive risk from the change itself: modules touched, dependency changes, schema migrations, public-contract modifications, ownership boundaries, sensitive-data access. A form that adds work to every change without changing where any change gets routed fails the test; a classifier that identifies which changes need specialized review can pass it.

> **Bounded production validation.** Production is part of verification, and an unacceptable first test for several classes of failure. Where failures are observable, reversible, and limited in blast radius, canary releases and defined rollback criteria verify behavior at a scale no review team matches. Production cannot be the primary oracle for unauthorized disclosure of personal data, destructive corruption, or irreversible migrations, because there the first observed failure is already unacceptable.

> **Compounding organizational memory.** A developer notices the agent keeps reaching for the wrong dependency, corrects it, ships the task. The next developer makes the same correction. A harness compounds only when that learning becomes durable: updated repository guidance, a regression case, a domain invariant, an automated architecture check, a retired instruction. Fixing the code resolves the incident. Changing the harness reduces the probability of the class. Memory is what prevents the rest of the harness from paying twice for the same lesson. Kief Morris describes this as moving humans from working *in the loop*, correcting individual outputs, to working *on the loop*, improving the system that produces them.

## Five implementations that frequently fail the test

The category is rarely invalid. What fails the test is the unmeasured, uncurated, universally applied implementation of it. A prompt library can demonstrate value under controlled comparison, a context file can reduce rework, and an architecture board can prevent a catastrophic boundary decision. These are the versions that usually cannot show it.

> **Detached prompt collections.** Prompts are useful; a collection of them is not a system. Most are weakly connected to specific failure modes, unversioned against outcomes, and rarely retired. A prompt joins the harness when it's embedded in a repeatable process and its effect can be measured.

> **Uncurated global context.** A repository file holding every standard, convention, and policy generates conflicting instructions, context pollution, and outdated guidance that keeps shaping new work. The objective is minimum sufficient context for the decision at hand, which requires curation, ownership, and a willingness to delete.

> **Non-enforcing governance artifacts.** A checklist earns its place when it changes routing, blocks an unsafe action, creates evidence, or triggers an automated control. One that produces an audit artifact without changing delivery behavior has done none of those. It may still be mandatory, and if it is, it answers to a different authority.

> **Recurring meetings for repeatable decisions.** Genuinely novel cross-system decisions deserve expert discussion. But when the same class of issue keeps requiring a meeting, the question is why that judgment was never converted into an architecture rule, a fitness function, or a routing policy. A standing meeting is often an organization paying rent on knowledge it already has.

> **Uniform specification and gating.** A full specification, architecture review, and staged rollout for a trivial visual adjustment isn't rigor. It's a tax, and people route around taxes, building a shadow delivery process nobody can observe or improve. Harness depth should follow risk, reversibility, and blast radius, not company size. A small team with low-stakes, reversible work should run very little harness. A small team crossing a payment, privacy, identity, or destructive-data boundary may need strict controls immediately.

## The human-review line sits at the control plane

The interesting disagreement is no longer whether a harness is needed. It's how far the harness should let humans move away from individual artifacts. I have set out the [layers of that harness](https://www.shiftharness.tech/quality-harness-engineering-the-emerging-stack-for/) at length elsewhere. The question here is which of them earn their keep once qualified attention is the thing being spent.

A pipeline gets very good at verifying routine behavior: tests pass, contracts stay compatible, errors stay flat, known invariants hold. It's much weaker at recognizing the significance of a structural decision nobody explicitly made. A feature can behave correctly while relocating an authorization boundary, creating a second source of truth, or setting a precedent future agents will copy. None of those trigger an immediate error.

> **On the loop for routine behavior. Explicit accountable approval for material system-boundary or control-plane changes.**

The trigger list is what makes that operational. Authorization or trust boundary. Data ownership or isolation boundary. Public contract or persistent schema. Irreversible side effect. Financial or regulated decision. Dependency direction or system-of-record ownership. And the one most teams miss: any harness rule, oracle, permission, or routing change that affects future releases. Repository agent instructions, agent permissions, CI bypass rules, test oracles, risk-tier classifiers, ownership metadata, deployment policies, security exceptions, deletion of a regression test.

A six-line change to an agent instruction file can be more consequential than a 600-line feature, because it shapes how the agent approaches and checks every change it touches from then on. Most review processes weight it by diff size and route it to nobody in particular.

Boundary movement can be flagged mechanically through dependency graphs, ownership metadata, public API modifications, and sensitive-path detection. It can't always be interpreted mechanically. The system can identify where judgment is needed. A human still owns the judgment.

![Three metal strips, each etched with a different export function above the same authorization check, identical on all three; only the centre strip carries an APPROVED stamp.](https://storage.ghost.io/c/73/3e/733efc05-c397-4cf1-b19b-527c1b07dee7/content/images/2026/09/image-4-1.png)

## The failure that has no strawman in it

Most illustrations of AI delivery risk describe something so obviously wrong that no competent team would ship it. Comfortable examples are part of why the real pattern keeps surviving.

I keep seeing the same failure, and it never looks like a failure while it's happening.

A team builds an AI-assisted customer-export feature. It has a specification. It has tests. It passes review. It works in production for months. Then a second team builds a different export capability. The coding agent reads the repository, finds the first implementation, and follows it as approved precedent, permission check included. That check was correct for the original feature's scope. It's subtly wrong for the new one. Now it exists in two places, and neither looks wrong read locally. A third implementation copies the same pattern, and with each repetition it looks more legitimate, because it's now more consistent with the codebase.

Every local control passes. The specification is satisfied. The tests reflect the implementation. Review finds the change consistent with existing code. Production error rates stay normal. The authorization gap only affects an unusual account configuration representing a small fraction of customers, which is why nobody sees it for two more quarters.

No individual stage is broken. The system's intent was never externalized strongly enough to resist the propagation of a locally reasonable precedent.

That's what the constraint looks like when it binds. It isn't sloppy code generation. It's verification proving conformance to a system whose intended boundary nobody preserved.

Three controls could have prevented or detected it, and none is a meeting: an authorization property derived independently of either implementation and exercised against the new export path, an architecture check enforcing centralized ownership of permission logic, and a repository rule plus regression case added the first time the duplication appeared.

## Subtracting a control safely is its own procedure

A thesis that promises subtraction owes you a method for it. Removing a control on intuition is the same failure as adding one on intuition, running in the other direction.

1. **Name the claimed risk.** What failure class does it prevent, detect, contain, or route?
2. **Measure current yield.** What did it change in the last two or three delivery windows?
3. **Replay known failures.** Does it detect real incidents from your own history?
4. **Identify overlapping coverage.** Which remaining controls cover the same class?
5. **Shadow before you stop blocking.** Keep the existing control enforcing while you log what a lighter replacement would have caught. Switching a blocking gate to advisory already removes its protection, so treat that switch as the removal step.
6. **Remove from a low-risk slice.** A repository, module, or tier with reversible impact.
7. **Monitor predefined guardrails.** Escapes, rework, verification queue time, release latency, owner interventions.
8. **Retire or restore.** Retire only if the slice actually exercised the failure class the control claims; if it didn't, the result is inconclusive and the control stays. Preserve a rollback path and record why.

The same procedure handles harness rot. Guidance written for a failure one model generation produced stays in place long after the models stopped producing it, consuming context and contradicting newer instructions. Rules need owners, evidence of continued value, and retirement criteria, or guidance ends up managed with less discipline than dependencies.

## Three questions worth instrumenting

None of these requires the authority to remove a control. Instrumentation is usually how that authority gets built.

> **How much change is verified without a human inspecting every artifact?** Measure the proportion of changes supported by credible automated evidence: specification-derived tests, domain invariants, contract checks, fitness functions, controlled production feedback. The goal is to reserve human attention for ambiguity, material risk, and structural judgment.

> **Which controls changed an outcome?** For each gate, ask what it blocked, corrected, rerouted, and what latency it added. Track gate yield, escape rate, false positives, and replay sensitivity, not pass rate alone. A gate that passes every change every time may be deterring failures in a low-incidence, high-impact domain, but deterrence should be demonstrated rather than assumed. Most **engineering efficiency metrics** never ask this, which is why the harness only ever grows.

> **How often does an escaped defect improve the system?** Call it escape-to-control conversion. The numerator is escaped defects for which, within a window you set, say 30 days, a durable control was merged and demonstrated through incident replay or seeded violation to detect, prevent, contain, or route that failure, including at least one variant beyond the original incident. Filing a ticket doesn't count, and neither does merging an artifact nobody tested against the incident. The control also needs an owner and a retirement condition.

Two things about the denominator. A defect that escaped despite an existing control stays eligible, because that is evidence the control is incomplete, misconfigured, bypassed, or no longer effective, and it should produce a strengthened invariant, a repaired routing rule, or the explicit retirement of a control that falsely claimed coverage. And do not require an incident review for eligibility, or the organization improves the metric by reviewing fewer incidents.

## What evidence would weaken the claim

The claim would weaken if, across multiple production organizations, AI-assisted change volume rose materially while total qualified verification attention, verification queue time, and reviewer utilization stayed flat or fell, and stability, escaped defects, rework, and time to restore did not deteriorate over a sufficiently long observation window.

The oracle argument would weaken if tests generated from the same requirement-and-implementation path detected specification-level misunderstandings as reliably as checks derived from an independent source of truth.

Neither result should be inferred from self-report or coarse merge metrics alone. The relevant evidence is whether additional change was absorbed without consuming more verification capacity or shifting cost into later failures and maintenance.

## The part that is hard to copy

Access to strong coding models isn't a durable advantage. Competitors buy the same models, the same agents, the same context windows. Whatever edge exists there closes on a purchase-order timeline.

What resists copying is a system that turns intent into checkable evidence, separates implementation from its definition of correctness, and converts meaningful failures into durable controls. Managed as a portfolio, that operating-model layer gets pruned about as often as it gets extended. Accumulated as a checklist, it only ever grows. That distinction is the lens I apply to this work, and I call it Shift Harness.

For a growing class of software work, generation is no longer the scarce capability. Trustworthy change is.

> **AI Transparency Notice:** This article and its accompanying images were created with the assistance of generative AI. The author directed the content, contributed the underlying ideas and analysis, and reviewed the final publication.

## Frequently Asked Questions

How do you decide whether an AI delivery control is worth keeping?▸

Where qualified verification attention is the binding constraint, a discretionary control earns its place only if it makes trustworthy change cheaper to verify. It has to reduce the qualified human attention needed for the same risk coverage, widen coverage at the same attention cost, or increase verified throughput without pushing the cost into escaped defects, delivery latency, or lost system understanding. Controls required by law, segregation of duties, or catastrophic-risk policy answer to another authority; the test can't remove them, only expose an expensive implementation.

Can AI-generated tests verify AI-generated code?▸

Not on their own. When one model interprets the requirement, writes the implementation, and generates the tests, the failure modes are correlated. A misunderstood requirement produces code implementing the wrong behavior, tests confirming that same wrong behavior, and an explanation justifying both with complete confidence. Coverage rises, test count rises, and the escaped-defect rate doesn't move. Writing the test first doesn't fix this, which is where the common advice falls short. Temporal order isn't independence. If the same model interpreted an ambiguous requirement, wrote the specification, generated the test, and wrote the implementation, the test existed first and all four artifacts still inherit the same misunderstanding. Independence is a property of provenance, at three levels. An independent truth source means correctness originates outside the implementation process, in a regulatory rule, a customer contract, a domain invariant, or an external API contract. Independent derivation means a different actor or process translates that truth into checks. Independent execution means the check runs outside the implementation's own assumptions and fixtures, for example a contract the provider verifies separately, or a replay of real production traffic against separately owned expected results. A property test counts only when its invariants come from the truth source rather than the code. A second model or an adversarial prompt still helps, but those are diversified reviewers rather than independent oracles, because different models share training data and similar blind spots. The design question is not who wrote the test. It is where the test's definition of correctness came from.

Why doesn't hiring more reviewers fix the AI code review bottleneck?▸

Because the scarce capacity is repository-specific judgment, not generic engineering hours. Distributing a diff across more people doesn't manufacture the context needed to spot a bad assumption, follow a structural implication, and take responsibility for a high-impact change. A senior architect's thirty minutes is not interchangeable with thirty minutes of generic review capacity. Margaret-Anne Storey's triple-debt model explains the mechanism. Technical debt lives in the code, cognitive debt lives in people as shared understanding erodes, and intent debt lives in externalized knowledge when goals, constraints, and rationale are never captured well enough for the next human or agent to use. Practice is heavily optimized for the first. AI-assisted delivery creates conditions in which the other two may grow faster, because implementation expands before shared understanding and externalized rationale are replenished. Work arrives syntactically clean and functional while fewer people can say why it's correct.

What should still require explicit human approval when AI writes most of the code?▸

Material system-boundary and control-plane changes. A pipeline gets very good at verifying routine behavior, so humans can work on the loop for that. It's much weaker at recognizing the significance of a structural decision nobody explicitly made, because a feature can behave correctly while relocating an authorization boundary or creating a second source of truth, and none of that triggers an immediate error. A practical trigger list: authorization or trust boundary, data ownership or isolation boundary, public contract or persistent schema, irreversible side effect, financial or regulated decision, dependency direction or system-of-record ownership. The one most teams miss is any harness rule, oracle, permission, or routing change that affects future releases, including repository agent instructions, agent permissions, CI bypass rules, test oracles, risk-tier classifiers, ownership metadata, and deletion of a regression test. A six-line change to an agent instruction file can be more consequential than a 600-line feature, because it shapes how the agent approaches and checks every change it touches from then on. Most review processes weight it by diff size and route it to nobody in particular.

How do you remove an engineering control without flying blind?▸

Treat removal as a measured procedure rather than a judgment call, because removing a control on intuition is the same failure as adding one on intuition. Name the claimed risk and the failure class it prevents, detects, contains, or routes. Measure what it actually changed in the last two or three delivery windows. Replay known failures from your own history against it. Identify which remaining controls cover the same class. Log what it would have caught while it keeps blocking, because turning it advisory is already a removal. Remove it from a low-risk slice with reversible impact. Monitor predefined guardrails, meaning escapes, rework, verification queue time, release latency, and owner interventions. Then retire or restore, preserving a rollback path and recording why. If the slice never exercised the failure class the control claims, the result is inconclusive and the control stays. The same procedure handles harness rot. Guidance written for a failure one model generation produced tends to stay in place long after the models stopped producing it, consuming context and contradicting newer instructions. Rules need owners, evidence of continued value, and retirement criteria, or guidance ends up managed with less discipline than dependencies.

How do you measure whether an AI engineering harness is working?▸

Measure verified change per unit of qualified human attention, calculated within risk tiers rather than as a single portfolio number. Setup cost and operating cost need separating too, because a fitness function can be expensive for eight weeks and nearly free for two years afterwards, and judging it inside the installation window rejects precisely the controls that amortize best. Three questions are worth instrumenting even without the authority to remove anything. How much change is verified without a human inspecting every artifact, measured as the proportion supported by credible automated evidence. Which controls actually changed an outcome, tracked through gate yield, escape rate, false positives, and replay sensitivity rather than pass rate alone, since a gate that passes every change every time may be deterring failures or may be doing nothing. And how often an escaped defect improves the system, meaning the share of escapes for which a durable control was merged within a defined window and demonstrated through incident replay or seeded violation to detect, prevent, contain, or route that failure, including at least one variant beyond the original incident. Filing a ticket doesn't count.