Research briefing · CMFC follow-up

When Models Fail to Repair the Same Mistake

Shared corruption and exact-repair failure across eight deployment routes.

Abstract

Giving several AI systems the same bad input might leave their average performance almost unchanged while making them fail together more often. In this finite synthetic study, eight deployment routes received either the same corruption variant or independent draws from the same pool. Average unsuccessful exact repair moved from 59.23% to 59.91%; the fraction of items where all eight routes failed exact repair moved from 4.59% to 6.98%. These are postconfirmatory findings, and the interval for the average change includes zero.

The implemented primary analysis also found an unsuccessful-repair dependence ratio above one. However, the study did not meet its full confirmatory-success rule: the implemented weighting differed from the protocol, and an omitted precollection support gate retrospectively fails. The result concerns how failures concentrate in this particular synthetic panel. It does not establish the frequency or cost of real-world incidents.

01 / Why this matters

An average can hide which cases fail together

Suppose a system sends a questionable record to several model routes for checking. A useful question is how often each route gets the task right. Another is whether their failures land on the same records. Two panels could have similar average accuracy yet leave very different numbers of cases without a successful response from any member.

Our original CMFC study separated shared task difficulty, residual dependence and common shocks. This follow-up narrows the intervention: when the corruption pool stays fixed, what changes if every route receives the same complete corruption variant rather than drawing independently?

The distinction matters because using several endpoints does not by itself establish useful redundancy. It also does not mean a shared input necessarily defeats that redundancy. The experiment measures a specific assignment contrast, rather than assuming either outcome.

02 / What was tested

The same pool, two ways of assigning errors

The study planned 1,800 item families across arithmetic, constraint, extraction and tool-call tasks. Code was excluded prospectively after calibration failure. Each retained family had a clean condition, a shared-corruption condition and an IID-corruption condition. “IID” means independent and identically distributed draws, with replacement, from the same pool of variants.

A route means a specific model/provider/deployment endpoint at the collection epoch. The eight routes should not be read as eight independent model families. The primary analysis retained 1,677 item families with complete support across all eight routes and all three conditions. Its dependence contrast compares the two corrupt conditions; the clean condition supplies a baseline for separate comparisons.

Figure 1 / The assignment changes, not the corruption pool

Shared corruption

Every route receives the same complete corruption variant.

Route 1A
Route 2A
Route 3A
Route 4A
Route 5A
Route 6A
Route 7A
Route 8A

IID corruption

Each route draws from the same pool, independently and with replacement.

Route 1A
Route 2C
Route 3B
Route 4D
Route 5A
Route 6B
Route 7D
Route 8C

Schematic only. Letters stand for complete corruption variants, not model identities or recorded assignments. Independent draws may coincide; the IID arm does not guarantee eight different errors. Both arms use the same pool. A separate clean condition supplies the baseline.

Responses had to flag the corruption, identify its target, supply the clean replacement, emit the designated repair object and return a terminal answer. The analysis kept observable non-detection, unsuccessful exact repair and an incorrect final answer separate. A route could answer correctly without emitting the required repair object. Conversely, flagging a problem did not establish that it was repaired.

“Exact repair” is therefore a strict output-contract test, not a judgment that every response missing that contract was useless. Nor does an emitted flag establish awareness or describe an internal cognitive stage.

03 / Average and tail

A small mean change, more all-eight failures

Figure 2 / Similar averages, a larger all-eight tail
Postconfirmatory summaries · 1,677 complete-support item families
IIDShared

Mean unsuccessful repair

Average fraction of routes failing exact repair per item

IID
59.23%
Shared
59.91%
0255075100%

All-eight unsuccessful repair

Fraction of item families where every route fails exact repair

IID
4.59%
Shared
6.98%
0255075100%

Both panels use the same zero-based percentage scale.

Mean change: +0.68 percentage points [−0.12, 1.45]. All-eight change: +2.39 percentage points [0.95, 3.82]. These are postconfirmatory, unadjusted percentile intervals, not the primary simultaneous analysis. The plots show observed rates, not deployment incident probabilities.

The average unsuccessful-repair rate rose by 0.68 percentage points, with an interval crossing zero. The all-eight count rose from 77 to 117 out of the same 1,677 item families. Its absolute increase was 2.39 percentage points. That divergence is the most accessible description of what changed: little movement in average failure, but more cases where no route emitted the exact repair.

Not every threshold moved upward. At the separate “at least half failed” threshold—four or more of eight, not a strict majority—the rate decreased slightly, from 69.65% to 68.99%. The all-eight result should not be generalized into a claim that every form of collective failure increased.

These mean and tail comparisons were postconfirmatory analyses. Their percentile intervals were unadjusted and do not inherit the simultaneous coverage of the primary analysis. An all-eight event also does not mean the routes are “fully correlated”; it is one observed outcome in the finite panel.

04 / The primary analysis

Repair dependence rose in the implemented analysis

Dependence inflation, or DI, compares the item-to-item variance of the panel’s failure fraction with the variance expected from its route marginal variances alone, leaving out cross-route covariance. This asks how concentrated failures are across items relative to that independence baseline. It is not simply another name for the failure rate.

The shared-to-IID DI ratio asks how that concentration measure changes between the two assignments. One means no change. The repair ratio was 1.129; the intervals for non-detection and terminal task failure both included one.

Figure 3 / The frozen implemented primary analysis
Simultaneous intervals · Full confirmatory-success rule not met
Non-detection1.038 [0.948, 1.137]
Non-detection: shared-to-IID DI ratio 1.038, simultaneous interval 0.948 to 1.137.

Interval includes one.

Unsuccessful exact repair1.129 [1.060, 1.203]
Unsuccessful exact repair: shared-to-IID DI ratio 1.129, simultaneous interval 1.060 to 1.203.

Interval above one; full study confirmation criteria not met.

Terminal task failure1.061 [0.992, 1.135]
Terminal task failure: shared-to-IID DI ratio 1.061, simultaneous interval 0.992 to 1.135.

Interval includes one.

0.901.001.101.201.25

Shared / IID dependence-inflation ratio · Dashed line = no change

9,999 stratified item-family bootstrap resamples. The repair interval lies above one, but the point estimator used equal item weights rather than the protocol’s fixed equal stratum weights. A required precollection calibration-support projection gate was omitted and retrospectively fails. The interval does not erase either deviation.

The qualifications change what can be claimed, not just how the result is described. Stratified resampling preserved the observed stratum counts, but that did not implement the protocol’s fixed equal stratum weights in the point estimate. Separately, the missing support projection was supposed to check that each endpoint-condition-outcome combination had enough expected failures and successes before collection. Reconstructing it afterward exposed failures.

The observed repair contrast remains part of the record. It should be presented alongside those departures rather than as a clean, fully confirmed study outcome.

05 / What the follow-ups resolve

Sensitivity checks preserve a direction, not every assumption

Named weighting, condition-specific support and missing-cell scenarios retained repair-dependence ratios above one. Those checks help establish that the reported direction was not unique to one retained specification. They remain postconfirmatory and do not retroactively satisfy the omitted precollection gate.

Under unrestricted binary completions of the missing outcomes, the reported outer envelope spans 0.989–1.311 and includes one. This envelope is not a confidence interval. It shows why the result cannot be called robust to arbitrary outcome-dependent missingness, even though the named scenarios preserved its direction.

The intervention also shares a complete corruption variant, which bundles features such as false value and target position. It does not isolate those features one by one. Independent assignment can produce accidental matches, and the bounded package does not contain the realized collision count and topology needed to characterize their effect on the nonlinear contrast.

06 / What to take from it

Measure the cases where no route repairs the error

The study gives a reason to inspect more than average correctness when evaluating a multi-route system. Marginal performance and all-member failure answer different questions. If a deployment depends on at least one route correcting an input, its evaluation needs to examine the joint outcomes directly rather than infer them from vendor counts or average scores.

It does not yet supply a deployment rule or identify the best way to diversify a fleet. A stronger confirmation would need the intended estimator, prospective support checks and evidence suited to the actual task. For this panel, the useful finding is narrower: shared corruption assignment changed the pattern of exact-repair failures in a way that the average alone did not capture.

07 / Paper and evidence

Read the methods alongside the result

This briefing is based on the retained D1 follow-up package, the reviewed paper and its numerical ledger. The package includes report-level results, sensitivity tables, analysis code, reviews and provenance receipts. Restricted prompts, responses, gold state and event databases are not included in the public release.

Download the reported headline results (CSV) or the machine-readable study summary (JSON). The download notes describe their scope. These are selected report-level aggregates, not a raw-data replication package.

The separate A2-F closeout concerns synthetic-fixture and API verification. It is not another empirical D1 analysis, and it does not complete the separate unfinished statistical-validation campaign. Nothing here claims independent regeneration of the historical model calls.

  1. Zinner and Beacon Bot. Shared Corruption Identity and Cross-Route Exact-Repair Failure Dependence: A Finite Synthetic Eight-Route Study. Public research working paper, September 2026. Not peer reviewed. Full references and retained-artifact citations are in the PDF.
  2. When Do Diverse Models Fail Together? Original Future Shock research briefing, August 2026.

Self-funded by Nicholas Zinner; no external funding or competing interests to declare. This is a public research working paper, not peer reviewed. Publication does not grant rights to restricted upstream artifacts.

AI news, analysis, and weekly deep dives. No hype.