When Do Diverse Models Fail Together? · Technical note
Methodology and public receipts
Populations, conditions, estimands, support gates, diagnostics, amendments, and disclosure boundaries for the CMFC common-shock study.
01 / Research question
How much does sharing the same recoverable instruction or evidence error increase fleet failure dependence relative to distributing severity-matched errors independently, while holding the underlying task and baseline difficulty fixed?
The broader retrospective layer contains 395 models on 27,842 exact common items. Its execution-side selection audit reported 3,253 support-qualified model pairs from 77,815 possible unordered pairs (4.18%). The public Stage 1 bundle does not include the event-level pair-support matrix needed to regenerate that selection count.
02 / Endpoint populations
The legacy population contains nine endpoints from seven developer families. The original ten-endpoint execution failed its frozen infrastructure-support gate; Qwen3-30B/Alibaba was excluded in whole before outcome access. The first nine-endpoint continuation also failed its unchanged gate. The reported execution preserves those terminal cells and reopens only retryable transport errors under an outcome-blind five-attempt, low-concurrency recovery amendment.
| Route ID | Provider | Model |
|---|---|---|
| openai-gpt41-nano-openai | OpenAI | GPT-4.1 Nano |
| openai-gpt41-mini-openai | OpenAI | GPT-4.1 Mini |
| google-gemini25-flash-lite-vertex-eu | Gemini 2.5 Flash-Lite | |
| anthropic-claude45-haiku-anthropic | Anthropic | Claude 4.5 Haiku |
| meta-llama31-8b-groq | Meta / Groq | Llama 3.1 8B |
| meta-llama33-70b-groq | Meta / Groq | Llama 3.3 70B |
| qwen3-8b-alibaba | Alibaba | Qwen3 8B |
| mistral-small32-24b-mistral-eu | Mistral | Mistral Small 3.2 24B |
| deepseek-v3-0324-crusoe-bf16 | DeepSeek / Crusoe | DeepSeek V3-0324 |
The contemporary population contains six endpoints from six developer families, frozen before legacy outcomes.
| Route ID | Provider | Model |
|---|---|---|
| deepseek-v4-flash-deepinfra-fp8 | DeepSeek / DeepInfra | DeepSeek V4 Flash |
| minimax-m3-minimax-fp8 | MiniMax | MiniMax M3 |
| moonshot-kimi-k27-inceptron-int4 | Moonshot / Inceptron | Kimi K2.7 |
| nvidia-nemotron35-coreweave-bf16 | NVIDIA / CoreWeave | Nemotron 3.5 |
| qwen38-27b-chutes-fp8 | Alibaba / Chutes | Qwen3.8 27B |
| zai-glm53-zai-fp8 | Zhipu AI | GLM 5.3 |
Legacy and contemporary populations are analyzed separately and are never pooled. Fallback routing is disabled for all endpoints.
The contemporary model and provider names are the frozen route labels recorded by the package. They are provenance claims about the requested routes, not independent validation of the underlying model weights.
03 / Six experimental conditions
All conditions retain the same correct answer and deterministic scorer. Models are told that reviewer notes and evidence are untrusted and may contain one error; the operative task is always to derive and return the correct answer.
| ID | Name | Definition |
|---|---|---|
| O | Original | Original task without reviewer proposal or evidence packet. |
| N | Neutral rewrite | Length- and format-matched answer-preserving rewrite with no proposed answer or corrupted assertion. Serves as manipulation check. |
| A_shared | Shared instruction anchor | Every endpoint receives the same SHA-selected plausible but wrong reviewer proposal, with an instruction to verify it. |
| A_independent | Independent instruction anchors | Endpoints receive different wrong proposals sampled without replacement from the same frozen variant pool. |
| E_shared | Shared evidence corruption | Every endpoint receives the same recoverable one-error variant of an otherwise identical evidence packet. |
| E_independent | Independent evidence corruptions | Endpoints receive different recoverable one-error variants sampled without replacement from the same frozen pool. |
04 / Estimand definitions
Dependence inflation (DI)
Plain language: DI measures how much the fleet's actual failure variance exceeds what would be expected if every endpoint failed independently at its own observed rate. A DI of 1.0 means independence; values above 1.0 mean failures are more clustered than independence predicts.
Formal: For condition c, let Σc be the sample covariance matrix of endpoint failure vectors with equal weights wm = 1/M. Fleet variance is Vc = w′Σcw. Independent-diagonal variance is Dc = Σm wm² Var(Yimc). Then DIc = Vc / Dc.
DI ratio (primary estimand)
Plain language: the DI ratio compares the dependence inflation when all endpoints receive the same error with the dependence inflation when each endpoint receives a different error of matched severity. Values above 1.0 mean sharing the same error increased dependence.
Formal: θE = log(DIE_shared / DIE_independent). Report exp(θ) as the ratio.
Absolute volatility ratio (AVR)
Plain language: AVR is the ratio of fleet standard deviations between shared and independent arms. It captures whether shared errors make fleet-level outcomes more variable.
Formal: AVR = √(Vshared / Vindependent).
Mean loss difference
Plain language: the difference in average fleet failure rate between shared and independent arms. Reports whether sharing the error changes how often the fleet is wrong on average, separate from whether it changes how clustered the errors are.
05 / Support quality tiers
| Tier | Threshold | Interpretation |
|---|---|---|
| confirmatory_quality | ≥ 90% overall matched support and ≥ 80% in every domain | Full primary inference |
| estimable_degraded | ≥ 500/800 overall and ≥ 60% per domain | Requires bounds; no deployment recommendation |
| non_estimable | Below degraded floor | Not analyzed as a primary result |
Both legacy primary contrasts are classified estimable_degraded, not confirmatory_quality. The evidence-corruption contrast has 665/800 (83.1%) matched support overall; two domains (arithmetic 71.9%, extraction 76.9%) fall below the 80% per-domain confirmatory floor. The instruction-anchor contrast has 682/800 (85.3%) overall.
Both contemporary contrasts achieve confirmatory_quality support (94/100 and 97/100), but their intervals include no change.
06 / Missingness and bounds
Infrastructure failures, ambiguous deliveries, route mismatches, and scorer failures are never recoded as model failures. Missing cells are excluded, and sharp binary-fill bounds cover mean loss and fleet tail events. Primary contrasts use exact matched support: an item contributes only if all endpoints completed both conditions in a pair.
Legacy evidence-corruption mean-loss bounds straddle zero: [−0.0140, +0.0115]. Legacy instruction-anchor bounds also straddle zero: [−0.0099, +0.0149]. Terminal error types in the legacy execution include ConnectError (1,825), HTTP 429 (312), ReadError (239), RemoteProtocolError (186), TimeoutError (31), and orphaned-started attempts (63).
07 / Diagnostic failures
Neutral manipulation check
The neutral manipulation check failed in both populations. The prespecified criterion requires O–N mean failure within ±0.03 and DIN/DIO within [0.90, 1.10].
- Legacy: O–N mean difference = −0.0398 (outside ±0.03 margin). DI ratio = 0.995 (within margin).
- Contemporary: O–N mean difference = −0.0333 (outside ±0.03 margin). DI ratio = 1.012 (within margin).
The N rewrite is systematically harder than O by approximately 3–4 percentage points. This limits total anchor-versus-neutral interpretation. It does not directly test the shared-versus-independent ratio, but it weakens confidence that the transformation pipeline changed only the intended feature.
Absent contemporary canaries
The contemporary population has zero canary observations across all six endpoints and all three epochs (start, midpoint, end). The prespecified drift-detection mechanism was entirely absent. No canary-based claim about route or behavior drift is available.
Endpoint degeneracies
- Legacy — Claude 4.5 Haiku: 100% failure rate under both A_shared (787/787) and A_independent (789/789). This endpoint is a zero-variance contributor to the primary anchor DI contrast, reducing the effective anchor fleet to eight endpoints (28 nondegenerate pairs).
- Contemporary — Nemotron 3.5: 97–100% failure rate across all conditions (97/99 even in the O condition). This endpoint is degenerate in both evidence conditions, reducing the effective evidence fleet to five endpoints (10 nondegenerate pairs).
Singleton-stratum correction
In the contemporary bootstrap, two strata contained a single item-family cluster each. A postcollection implementation correction treats these as self-representing certainty strata. This correction was made after legacy results existed and is explicitly not described as outcome-blind. It changes bootstrap execution but no contemporary point estimate, item, verdict, stratum, or estimand.
08 / Summarized amendment register
Twenty named protocol amendments were made between contract freeze and final analysis. Most are claimed outcome-blind at their point of application; the singleton-stratum correction explicitly is not. The amendment chain is self-attested through frozen artifacts and SHA-256 hashes, not independently witnessed.
| Amendment | Timing | Summary |
|---|---|---|
| Simulation-only 1 | Pre-pilot | Fix exchangeability in null simulator; add 800-item candidate size. |
| Simulation-only 2 | Pre-pilot | Replace raw type-I cutoff with exact binomial Monte Carlo interval. |
| Pilot infrastructure 3 | Pre-pilot outcome | Irreversible quarantine → restart from zero on same frozen cells. |
| Pilot infrastructure 4 | Mid-pilot | Exclude ambiguous/503 cells as missing under frozen support floors. |
| Pilot support 5 | Mid-pilot | Lower pilot endpoint-condition floor from 97% to 95%. |
| Pilot active-time 6 | Mid-pilot | Exclude governance pauses from the eight-hour active-time ceiling. |
| Replacement pilot | Pre-confirmation | New pilot-v9 execution after first pilot failed its gate. |
| Pilot budget | Pre-confirmation | Increase pilot module cap from $2 to $2.75. |
| Analysis stratum cluster | Pre-confirmation | Freeze cluster definitions and stratum weights. |
| Domain tier adjustment | Pre-confirmation | Adjust construction-tier mixture for difficulty band. |
| Output-contract remediation | Pre-confirmation | Fix output-contract parsing for edge cases. |
| Provider route restart | Mid-confirmation | Restart after provider route returned 503 error envelope. |
| Arithmetic mixture adjustment | Mid-confirmation | Adjust arithmetic domain mixture after pilot gate results. |
| Error envelope restart | Mid-confirmation | Restart scheduler after persistent provider error envelopes. |
| Confirmation throughput restart | Mid-confirmation | Restart after scheduler throughput failure. |
| Confirmation streaming restart | Mid-confirmation | Continuously refilled scheduler and clean v9 restart. |
| Contemporary streaming plan | Pre-contemporary | Zero-call contemporary plan refresh. |
| Confirmation activation + budget | Pre-confirmation | Freeze realistic execution ceiling outcome-blind. |
| Transport recovery | Post-confirmation | Outcome-blind five-attempt low-concurrency recovery of retryable transport errors. |
| Singleton stratum correction | Post-legacy analysis | Self-representing certainty strata for singleton clusters in contemporary bootstrap. Not outcome-blind. |
09 / Evidence register
| Status | Artifact | Description |
|---|---|---|
| Included | ERRATA.md | Interpretation corrections for preserved upstream files |
| Included | base-study-report.md | 395-model retrospective + 10-endpoint prospective results |
| Included | common-shock-report.md | Primary contrasts, condition decomposition, budget |
| Included | common-shock-missingness.md | Missingness bounds by condition and contrast |
| Included | common-shock-contract.md | Study plan v2 with all frozen amendments |
| Included | legacy-summary.json | Machine-readable legacy population results |
| Included | contemporary-summary.json | Machine-readable contemporary population results |
| Included | pilot-gate.json | Pilot pass/fail diagnostics |
| Referenced | Local evidence estate | Raw event ledgers, matrices, bootstrap outputs, source manifests, checksums |
| Referenced | Collection and analysis code | Hash-described from workspace; some early patch state unavailable |
| Withheld | Restricted Phase-1 source bytes | Pending redistribution permission |
| Withheld | Raw provider payloads | Credentials, response bodies, session metadata |
| Withheld | Run databases | Internal operational state |
10 / Stage 1 reproducibility
The first public evidence bundle is a Stage 1 release. Stage 1 means report-level inspection and receipt reconstruction are supported. It does not mean full independent reproducibility: the historical model calls cannot be replayed, some early collection-time patch state is unavailable, and restricted Phase-1 source bytes are excluded pending redistribution permission.
The delivered package establishes internal consistency through SHA-256 checksums, completion receipts, and a manifest. It does not establish bit-exact external reproducibility. The study contract, estimands, cluster definitions, bootstrap parameters, and sample size were frozen before model calls, with simulation-validated operating characteristics (type I error, coverage, power) preceding any data collection.
11 / Inference
Primary common-shock intervals use a 9,999-resample studentized maximum-absolute-statistic bootstrap, clustering by item family within frozen domain and predicted-difficulty strata. The joint coverage target is 95% simultaneous across both primary contrasts (θA and θE).
The bootstrap was simulation-validated before any pilot call against targets of ≥ 90% power for a DI ratio of 1.25 and ≥ 80% power for 1.20, with familywise two-sided α 0.05. Simulation selected 800 confirmatory items. It did not change using observed effect sizes.
If DI or its ratio is undefined (e.g. zero-variance endpoint), the condition is reported as floor/ceiling saturated with mean and tail loss only. No post-outcome regularization is applied.