When Do Diverse Models Fail Together? · Technical note

Methodology and public receipts

Populations, conditions, estimands, support gates, diagnostics, amendments, and disclosure boundaries for the CMFC common-shock study.

01 / Research question

How much does sharing the same recoverable instruction or evidence error increase fleet failure dependence relative to distributing severity-matched errors independently, while holding the underlying task and baseline difficulty fixed?

The broader retrospective layer contains 395 models on 27,842 exact common items. Its execution-side selection audit reported 3,253 support-qualified model pairs from 77,815 possible unordered pairs (4.18%). The public Stage 1 bundle does not include the event-level pair-support matrix needed to regenerate that selection count.

02 / Endpoint populations

The legacy population contains nine endpoints from seven developer families. The original ten-endpoint execution failed its frozen infrastructure-support gate; Qwen3-30B/Alibaba was excluded in whole before outcome access. The first nine-endpoint continuation also failed its unchanged gate. The reported execution preserves those terminal cells and reopens only retryable transport errors under an outcome-blind five-attempt, low-concurrency recovery amendment.

Legacy nine-endpoint population
Route IDProviderModel
openai-gpt41-nano-openaiOpenAIGPT-4.1 Nano
openai-gpt41-mini-openaiOpenAIGPT-4.1 Mini
google-gemini25-flash-lite-vertex-euGoogleGemini 2.5 Flash-Lite
anthropic-claude45-haiku-anthropicAnthropicClaude 4.5 Haiku
meta-llama31-8b-groqMeta / GroqLlama 3.1 8B
meta-llama33-70b-groqMeta / GroqLlama 3.3 70B
qwen3-8b-alibabaAlibabaQwen3 8B
mistral-small32-24b-mistral-euMistralMistral Small 3.2 24B
deepseek-v3-0324-crusoe-bf16DeepSeek / CrusoeDeepSeek V3-0324

The contemporary population contains six endpoints from six developer families, frozen before legacy outcomes.

Contemporary six-endpoint population
Route IDProviderModel
deepseek-v4-flash-deepinfra-fp8DeepSeek / DeepInfraDeepSeek V4 Flash
minimax-m3-minimax-fp8MiniMaxMiniMax M3
moonshot-kimi-k27-inceptron-int4Moonshot / InceptronKimi K2.7
nvidia-nemotron35-coreweave-bf16NVIDIA / CoreWeaveNemotron 3.5
qwen38-27b-chutes-fp8Alibaba / ChutesQwen3.8 27B
zai-glm53-zai-fp8Zhipu AIGLM 5.3

Legacy and contemporary populations are analyzed separately and are never pooled. Fallback routing is disabled for all endpoints.

The contemporary model and provider names are the frozen route labels recorded by the package. They are provenance claims about the requested routes, not independent validation of the underlying model weights.

03 / Six experimental conditions

All conditions retain the same correct answer and deterministic scorer. Models are told that reviewer notes and evidence are untrusted and may contain one error; the operative task is always to derive and return the correct answer.

Condition definitions
IDNameDefinition
OOriginalOriginal task without reviewer proposal or evidence packet.
NNeutral rewriteLength- and format-matched answer-preserving rewrite with no proposed answer or corrupted assertion. Serves as manipulation check.
A_sharedShared instruction anchorEvery endpoint receives the same SHA-selected plausible but wrong reviewer proposal, with an instruction to verify it.
A_independentIndependent instruction anchorsEndpoints receive different wrong proposals sampled without replacement from the same frozen variant pool.
E_sharedShared evidence corruptionEvery endpoint receives the same recoverable one-error variant of an otherwise identical evidence packet.
E_independentIndependent evidence corruptionsEndpoints receive different recoverable one-error variants sampled without replacement from the same frozen pool.

04 / Estimand definitions

Dependence inflation (DI)

Plain language: DI measures how much the fleet's actual failure variance exceeds what would be expected if every endpoint failed independently at its own observed rate. A DI of 1.0 means independence; values above 1.0 mean failures are more clustered than independence predicts.

Formal: For condition c, let Σc be the sample covariance matrix of endpoint failure vectors with equal weights wm = 1/M. Fleet variance is Vc = w′Σcw. Independent-diagonal variance is Dc = Σm wm² Var(Yimc). Then DIc = Vc / Dc.

DI ratio (primary estimand)

Plain language: the DI ratio compares the dependence inflation when all endpoints receive the same error with the dependence inflation when each endpoint receives a different error of matched severity. Values above 1.0 mean sharing the same error increased dependence.

Formal: θE = log(DIE_shared / DIE_independent). Report exp(θ) as the ratio.

Absolute volatility ratio (AVR)

Plain language: AVR is the ratio of fleet standard deviations between shared and independent arms. It captures whether shared errors make fleet-level outcomes more variable.

Formal: AVR = √(Vshared / Vindependent).

Mean loss difference

Plain language: the difference in average fleet failure rate between shared and independent arms. Reports whether sharing the error changes how often the fleet is wrong on average, separate from whether it changes how clustered the errors are.

05 / Support quality tiers

Support classification
TierThresholdInterpretation
confirmatory_quality≥ 90% overall matched support and ≥ 80% in every domainFull primary inference
estimable_degraded≥ 500/800 overall and ≥ 60% per domainRequires bounds; no deployment recommendation
non_estimableBelow degraded floorNot analyzed as a primary result

Both legacy primary contrasts are classified estimable_degraded, not confirmatory_quality. The evidence-corruption contrast has 665/800 (83.1%) matched support overall; two domains (arithmetic 71.9%, extraction 76.9%) fall below the 80% per-domain confirmatory floor. The instruction-anchor contrast has 682/800 (85.3%) overall.

Both contemporary contrasts achieve confirmatory_quality support (94/100 and 97/100), but their intervals include no change.

06 / Missingness and bounds

Infrastructure failures, ambiguous deliveries, route mismatches, and scorer failures are never recoded as model failures. Missing cells are excluded, and sharp binary-fill bounds cover mean loss and fleet tail events. Primary contrasts use exact matched support: an item contributes only if all endpoints completed both conditions in a pair.

Legacy evidence-corruption mean-loss bounds straddle zero: [−0.0140, +0.0115]. Legacy instruction-anchor bounds also straddle zero: [−0.0099, +0.0149]. Terminal error types in the legacy execution include ConnectError (1,825), HTTP 429 (312), ReadError (239), RemoteProtocolError (186), TimeoutError (31), and orphaned-started attempts (63).

07 / Diagnostic failures

Neutral manipulation check

The neutral manipulation check failed in both populations. The prespecified criterion requires O–N mean failure within ±0.03 and DIN/DIO within [0.90, 1.10].

The N rewrite is systematically harder than O by approximately 3–4 percentage points. This limits total anchor-versus-neutral interpretation. It does not directly test the shared-versus-independent ratio, but it weakens confidence that the transformation pipeline changed only the intended feature.

Absent contemporary canaries

The contemporary population has zero canary observations across all six endpoints and all three epochs (start, midpoint, end). The prespecified drift-detection mechanism was entirely absent. No canary-based claim about route or behavior drift is available.

Endpoint degeneracies

Singleton-stratum correction

In the contemporary bootstrap, two strata contained a single item-family cluster each. A postcollection implementation correction treats these as self-representing certainty strata. This correction was made after legacy results existed and is explicitly not described as outcome-blind. It changes bootstrap execution but no contemporary point estimate, item, verdict, stratum, or estimand.

08 / Summarized amendment register

Twenty named protocol amendments were made between contract freeze and final analysis. Most are claimed outcome-blind at their point of application; the singleton-stratum correction explicitly is not. The amendment chain is self-attested through frozen artifacts and SHA-256 hashes, not independently witnessed.

Summarized amendment chronology
AmendmentTimingSummary
Simulation-only 1Pre-pilotFix exchangeability in null simulator; add 800-item candidate size.
Simulation-only 2Pre-pilotReplace raw type-I cutoff with exact binomial Monte Carlo interval.
Pilot infrastructure 3Pre-pilot outcomeIrreversible quarantine → restart from zero on same frozen cells.
Pilot infrastructure 4Mid-pilotExclude ambiguous/503 cells as missing under frozen support floors.
Pilot support 5Mid-pilotLower pilot endpoint-condition floor from 97% to 95%.
Pilot active-time 6Mid-pilotExclude governance pauses from the eight-hour active-time ceiling.
Replacement pilotPre-confirmationNew pilot-v9 execution after first pilot failed its gate.
Pilot budgetPre-confirmationIncrease pilot module cap from $2 to $2.75.
Analysis stratum clusterPre-confirmationFreeze cluster definitions and stratum weights.
Domain tier adjustmentPre-confirmationAdjust construction-tier mixture for difficulty band.
Output-contract remediationPre-confirmationFix output-contract parsing for edge cases.
Provider route restartMid-confirmationRestart after provider route returned 503 error envelope.
Arithmetic mixture adjustmentMid-confirmationAdjust arithmetic domain mixture after pilot gate results.
Error envelope restartMid-confirmationRestart scheduler after persistent provider error envelopes.
Confirmation throughput restartMid-confirmationRestart after scheduler throughput failure.
Confirmation streaming restartMid-confirmationContinuously refilled scheduler and clean v9 restart.
Contemporary streaming planPre-contemporaryZero-call contemporary plan refresh.
Confirmation activation + budgetPre-confirmationFreeze realistic execution ceiling outcome-blind.
Transport recoveryPost-confirmationOutcome-blind five-attempt low-concurrency recovery of retryable transport errors.
Singleton stratum correctionPost-legacy analysisSelf-representing certainty strata for singleton clusters in contemporary bootstrap. Not outcome-blind.

09 / Evidence register

Included / referenced / withheld evidence
StatusArtifactDescription
IncludedERRATA.mdInterpretation corrections for preserved upstream files
Includedbase-study-report.md395-model retrospective + 10-endpoint prospective results
Includedcommon-shock-report.mdPrimary contrasts, condition decomposition, budget
Includedcommon-shock-missingness.mdMissingness bounds by condition and contrast
Includedcommon-shock-contract.mdStudy plan v2 with all frozen amendments
Includedlegacy-summary.jsonMachine-readable legacy population results
Includedcontemporary-summary.jsonMachine-readable contemporary population results
Includedpilot-gate.jsonPilot pass/fail diagnostics
ReferencedLocal evidence estateRaw event ledgers, matrices, bootstrap outputs, source manifests, checksums
ReferencedCollection and analysis codeHash-described from workspace; some early patch state unavailable
WithheldRestricted Phase-1 source bytesPending redistribution permission
WithheldRaw provider payloadsCredentials, response bodies, session metadata
WithheldRun databasesInternal operational state

10 / Stage 1 reproducibility

The first public evidence bundle is a Stage 1 release. Stage 1 means report-level inspection and receipt reconstruction are supported. It does not mean full independent reproducibility: the historical model calls cannot be replayed, some early collection-time patch state is unavailable, and restricted Phase-1 source bytes are excluded pending redistribution permission.

The delivered package establishes internal consistency through SHA-256 checksums, completion receipts, and a manifest. It does not establish bit-exact external reproducibility. The study contract, estimands, cluster definitions, bootstrap parameters, and sample size were frozen before model calls, with simulation-validated operating characteristics (type I error, coverage, power) preceding any data collection.

11 / Inference

Primary common-shock intervals use a 9,999-resample studentized maximum-absolute-statistic bootstrap, clustering by item family within frozen domain and predicted-difficulty strata. The joint coverage target is 95% simultaneous across both primary contrasts (θA and θE).

The bootstrap was simulation-validated before any pilot call against targets of ≥ 90% power for a DI ratio of 1.25 and ≥ 80% power for 1.20, with familywise two-sided α 0.05. Simulation selected 800 confirmatory items. It did not change using observed effect sizes.

If DI or its ratio is undefined (e.g. zero-variance endpoint), the condition is reported as floor/ceiling saturated with mean and tail loss only. No post-outcome regularization is applied.

← Return to main briefing

AI news, analysis, and weekly deep dives. No hype.