Research briefing · Reliability measurement
When Do Diverse Models Fail Together?
Difficulty, Shared Evidence, and Common Shocks Across Language-Model Fleets
Abstract
Across 27,842 exact common retrospective items, median raw chance-adjusted agreement (CAPA) was 0.498, but median residual correlation after a two-parameter item-response adjustment was −0.018. Most of the broad co-failure signal was therefore consistent with shared item difficulty rather than a large average residual model common cause. High raw agreement alone is not evidence of a hidden model monoculture.
A prospective ten-endpoint panel retained small positive cross-family residual dependence (β 0.068 [0.039, 0.093]). Empirical failure clusters showed almost no alignment with named developer families: adjusted Rand index 0.012 retrospectively and −0.071 prospectively. The retrospective comparison covered 127 models in six eligible multi-model families; these descriptive results do not show that vendor identity is irrelevant, and a vendor count is not an independence count. The directional prediction that evidence-path diversification would outperform a fixed endpoint substitution was also rejected in both confirmations (Δ −0.460 [−0.589, −0.327] and −0.150 [−0.182, −0.119]). Those substitutions changed an endpoint bundle, not vendor identity alone.
Recoverable common shocks produced one modest legacy evidence signal at estimable degraded support; both contemporary intervals included no change, and neither instruction-anchor contrast excluded the null. Separately, a severe semantic instruction-conflict stress amplified joint failure 4.337-fold [3.358, 5.885]. That stress demonstrates a distinct synchronization mechanism, not its general prevalence in deployment.
01 / Three failure channels
Co-failure is not one phenomenon
A fleet of language models is not automatically a fleet of independent risks. Models may share training material, evaluation difficulty, retrieval systems, prompts, providers, or upstream evidence. Conversely, raw agreement on hard items can be large even when residual failures are nearly independent. The deployment question is conditional: how much independence remains after difficulty is controlled, and how much is lost when models receive a common erroneous input?
This study keeps three co-failure channels separate rather than treating model correlation as a single trait.
Two measurements do different jobs here. CAPA, chance-adjusted probabilistic agreement, asks whether two models fail together more often than their individual failure rates would suggest. A value near zero means their overlap is not larger than that chance baseline; a larger positive value means the overlap is stronger. That makes CAPA useful for describing a fleet. It does not tell us why the overlap exists. If an item is hard for nearly everyone, many models can fail together without sharing one hidden defect.
The 2PL adjustment tries to remove that ordinary difficulty channel before asking what remains. A two-parameter logistic item-response model gives each item a difficulty term and a discrimination term, then estimates how each model responds across that landscape. The residual correlation is the pairwise co-failure left after those item properties are accounted for. It is not a causal mechanism or a proof of independence. It is a narrower statement: once this model has absorbed the measured difficulty structure, the average residual pairwise signal is small.
That distinction is why the raw and adjusted findings can look so different without contradicting each other. The retrospective panel can show substantial chance-adjusted agreement while its median residual correlation sits near zero. One number describes what happened on the same questions. The other asks how much of that pattern survives a particular difficulty model. Neither says what will happen when a common retriever, prompt, or evidence packet injects the same mistake into every route.
A single correlation number can collapse these channels. This study separates them with distinct populations, designs, and estimands.
02 / Study map
From observation to controlled intervention
The design progresses through five stages, each with its own population, design, and estimand. The stages are analyzed separately and never pooled.
Each stage uses a different population and design. They are not pooled. The progression narrows from observation to controlled intervention.
The retrospective panel (395 models, 27,842 items) measures raw agreement. The prospective panel (10 fixed endpoints, 1,500 fresh items) measures residual dependence after difficulty adjustment. Evidence-path and endpoint contrasts test whether source diversity or endpoint diversity produces more decorrelation. The adversarial battery measures worst-case synchronization under semantic conflict. The common-shock experiment controls shared errors to estimate their causal effect on fleet dependence.
The five stages are not successive attempts to estimate one master correlation. Each removes a different ambiguity. The retrospective panel supplies breadth across hundreds of models but inherits uneven historical coverage. The prospective panel trades that breadth for ten fixed endpoints answering the same 1,500 new items. The source and endpoint contrasts then alter one part of the evidence path at a time. The adversarial and architecture batteries probe stronger mechanisms. Only the final common-shock experiment directly assigns shared versus independent wrong inputs while keeping the item, error burden, endpoint fleet, and scoring rule fixed.
Coverage matters in the broad panel. The retrospective execution's selection audit reported 3,253 support-qualified model pairs from 77,815 possible unordered pairs, about 4.18 percent. That is a large comparison set, but it is not a complete all-by-all fleet census. Pairwise results apply to routes with enough overlapping valid outcomes to clear the support rule. The prospective stages exist partly because same-item collection removes that historical-overlap problem. Results from those panels are reported beside the retrospective picture, not pooled into it. The Stage 1 bundle does not carry the event-level pair-support matrix, so a public reader cannot independently regenerate the 3,253 count from this release.
The later panels also differ in endpoint roster, provider route, collection epoch, and item bank. That separation is deliberate. Reusing the same calls would make the studies look consistent without showing that the result survives a new population. A separate panel can provide stronger evidence when its estimate is precise. Here the 100-item contemporary panel is too imprecise to do that for the recoverable evidence shock, so it remains a new measurement rather than a confirmation label.
03 / Raw vs. adjusted dependence
Most raw co-failure is shared difficulty
In the retrospective panel, raw median chance-adjusted agreement (CAPA) was 0.498 [0.494, 0.503]. But the median 2PL residual correlation after difficulty adjustment was −0.018. Most coincidence in this population is therefore consistent with shared item difficulty rather than a large average residual common-cause signal.
This result extends rather than ignores the earlier difficulty checks. Kim et al. identified unequal question difficulty as a limitation of wrong-answer agreement. Goel et al. stratified CAPA by coarse question-hardness bins and reported the capability–similarity trend as mostly stable. The present analysis instead models item difficulty and discrimination at item level, then measures pairwise residual dependence. The near-zero retrospective median shows that those methods need not yield the same interpretation; differences in estimand, population, pair support, and adjustment resolution remain plausible explanations.
In the prospective ten-endpoint panel, 1,495 of 1,500 items had exact common support. Cross-family residual dependence was small but positive: β 0.068 [0.039, 0.093]. The hardest retrospective decile was ceiling-saturated at 95.1% pooled failure; it did not identify tail independence or a parametric tail copula.
04 / Diversity contrasts
Source diversity is not endpoint diversity
The directional prediction that evidence-path diversification would exceed fixed endpoint substitution was rejected. In the 150-country primary confirmation, the endpoint substitution produced a larger decorrelation (Δ −0.460 [−0.589, −0.327]) than the evidence-path contrast. The 1,200-city secondary extension showed the same direction (Δ −0.150 [−0.182, −0.119]).
The endpoint contrast jointly changes model identity, family, size, provider, and precision. It is not a causal family-only effect. Regardless of the mechanism, the practical result is that swapping the model endpoint produced more decorrelation than swapping the evidence source.
"Evidence diversity" is narrower here than the phrase may sound. The country and city tasks use paired source routes that agree on overlapping facts. Changing the route can alter retrieval, presentation, or source-specific noise, but the design does not prove that the two routes are epistemically independent or grounded in disjoint upstream records. The result therefore compares the observed decorrelation from those evidence paths with the observed decorrelation from changing endpoints. It does not establish a general hierarchy in which model diversity always beats source diversity.
The endpoint substitution is broad in a different way. It jointly changes the named model, developer family, parameter scale, inference provider, and precision. That bundle produced more decorrelation in both confirmations, but the experiment cannot assign the gain to one member of the bundle. This matters operationally: a team cannot take the coefficient and conclude that switching providers, families, or parameter counts alone will recreate it. The supported claim is about these exact endpoint substitutions on these tasks.
05 / Six experimental conditions
Shared versus independent errors
The common-shock experiment crossed 800 fresh items (five balanced domains) with six conditions on a nine-endpoint legacy fleet. A separately frozen 100-item contemporary panel used six endpoints. All conditions retain the same correct answer and deterministic scorer.
- O — Original task without added proposal or evidence packet.
- N — Answer-preserving neutral rewrite (manipulation check).
- A_shared / A_independent — Same or independently assigned wrong reviewer proposal; models are told it is untrusted.
- E_shared / E_independent — Same or independently assigned one-error evidence corruption; the error is recoverable.
Dependence inflation, or DI, turns pairwise co-failure into a fleet-level variance quantity. For one condition, the numerator is the observed variance of the equally weighted fleet failure rate. The denominator is the variance expected from the endpoints' own Bernoulli failure rates if off-diagonal covariance were absent. A DI near 1 means the observed fleet variance is close to that independent-diagonal reference. A value above 1 means positive co-movement inflates it. A value below 1 means the covariance contribution is negative. DI is still a summary of binary outcomes, not an explanation of architecture, training lineage, or provider behavior.
The primary comparison is not a condition's DI in isolation. It is the ratio of DI under a shared corruption to DI under independently assigned corruptions of the same type. Items enter a contrast only when every endpoint has valid outcomes in both arms, so the numerator and denominator use exact matched support. This paired design holds the underlying question and corruption severity fixed while changing the identity of the bad input. It is the cleanest available test of whether common input identity adds dependence beyond the marginal damage caused by receiving a wrong proposal or evidence packet at all.
Uncertainty is also paired. The analysis resamples item-family clusters within frozen domain and predicted-difficulty strata, rather than pretending every rewritten variant is an independent observation. A 9,999-resample studentized maximum-statistic procedure produces simultaneous intervals for the anchor and evidence contrasts. In plain language, the intervals account for related items and for asking two primary common-shock questions together. Mean loss and absolute fleet volatility remain separate estimands because the same shock can change average error, covariance, and tail behavior by different amounts.
The primary estimands are the DI (dependence inflation) ratios comparing shared versus independent arms. Mean loss, absolute fleet volatility (AVR), and tail events are reported separately.
06 / Common-shock results
One degraded-support signal; three intervals include no change
Scroll horizontally for all columns →
| Population | Shock | Matched | Support | DI ratio | 95% CI | Δ mean | AVR |
|---|---|---|---|---|---|---|---|
| Legacy | Evidence corruption | 665 / 800 | estimable degraded | 1.095 | [1.016, 1.181] | +0.002 | 1.060 |
| Legacy | Instruction anchor | 682 / 800 | estimable degraded | 0.979 | [0.925, 1.036] | +0.003 | 0.989 |
| Contemporary | Evidence corruption | 97 / 100 | confirmatory quality | 1.007 | [0.888, 1.142] | −0.007 | 1.006 |
| Contemporary | Instruction anchor | 94 / 100 | confirmatory quality | 1.033 | [0.925, 1.153] | −0.016 | 1.030 |
The support label is part of the result, not a footnote to it. The two legacy contrasts were estimable, but both were classified as estimable degraded under the frozen support rules. The evidence interval barely clears 1.0, while the instruction-anchor interval includes 1.0. The contemporary evidence interval is wide enough to include no effect and an effect the size of the legacy estimate. It therefore neither precisely confirms nor rejects the recoverable evidence-shock effect. Calling it a failed replication would overstate the negative evidence; calling it confirmation would overstate the positive evidence.
Missing outcomes are not automatically model failures and were not recoded as such. That protects the scientific outcome from transport errors, route mismatches, ambiguous deliveries, and scorer failures. It also leaves a sensitivity problem when missingness is concentrated by endpoint or condition. The published sharp bounds show that the legacy mean-loss difference can change sign under extreme binary fills. They do not supply an equivalent sharp bound for the DI ratio. The safest reading is consequently narrow: one legacy dependence interval excludes no change on matched support, the support is degraded, and the delivered package does not establish how the dependence estimate behaves under every completion of the missing cells.
The only null-excluding result is the legacy shared-evidence contrast: DI ratio 1.095 [1.016, 1.181] with 665/800 matched items at estimable degraded support. The confidence interval lower bound is only 0.016 above 1.0 on the ratio scale. Missingness bounds show the mean-loss difference could flip sign. This is a modest, fragile signal at degraded support—not a confirmed effect.
The contemporary shared-evidence DI ratio was 1.007 [0.888, 1.142] at confirmatory quality support. This interval includes no change. Neither instruction-anchor contrast in either population excluded the null.
07 / Condition-level decomposition
What each condition looks like
Scroll horizontally for all columns →
| Condition | Items | Mean loss | DI | n_eff | Maj. fail | All fail |
|---|---|---|---|---|---|---|
| O | 736 | 0.446 | 3.688 | 2.431 | 0.409 | 0.037 |
| N | 733 | 0.487 | 3.633 | 2.471 | 0.447 | 0.060 |
| A_indep | 723 | 0.534 | 2.602 | 3.012 | 0.492 | 0.018 |
| A_shared | 733 | 0.536 | 2.557 | 3.067 | 0.499 | 0.015 |
| E_indep | 723 | 0.538 | 2.358 | 3.559 | 0.484 | 0.018 |
| E_shared | 717 | 0.538 | 2.533 | 3.349 | 0.492 | 0.018 |
Treatment conditions (A and E) increase mean loss from the 0.45 range to the 0.53–0.54 range while reducing absolute DI. The primary shared-versus-independent contrasts use paired within-item comparisons, not the absolute DI levels shown here.
The lower absolute DI values in the treated conditions are not evidence that misleading inputs make the fleet safer. DI divides total fleet variance by a marginal-variance reference. The anchor and evidence treatments push mean failure toward the middle of the Bernoulli range, increasing that denominator more than the observed off-diagonal covariance increases. Average loss can rise while absolute DI falls. That is exactly why the paper keeps condition-level loss, absolute DI, and the paired shared-versus-independent DI ratios in separate columns instead of compressing them into one reliability score.
08 / Mechanism differences
Do not visually pool incomparable mechanisms
A harmless semantic instruction-conflict attack amplified median pairwise joint failure 4.337-fold [3.358, 5.885]. Token-suffix attacks produced 0.826 [0.642, 1.086]. In the matched architecture battery, parallel proposals retained effective diversity of 4.533 while chains compressed to 1.580.
These adversarial and architectural results measure different mechanisms from the controlled common shocks. The adversarial instruction conflict changes the operative task. The common shocks are recoverable, severity-matched, one-error corruptions that preserve the correct answer. Their magnitudes should not be placed on the same visual scale.
A coherent wrong directive changes the operative task.
Adversarial transfer battery; not a common-shock DI ratio.
The same recoverable one-error evidence variant reaches every endpoint.
DI ratio; estimable_degraded support, 665/800 matched.
The same untrusted wrong proposal reaches every endpoint.
DI ratio; estimable_degraded support, interval includes unity.
The recoverable evidence intervention is repeated on six newer routes.
DI ratio; confirmatory-quality support, inconclusive interval.
The cards deliberately avoid a shared bar axis. The adversarial result is joint-failure amplification under semantic instruction conflict. The common-shock results are DI ratios under recoverable, severity-matched errors. Comparing their lengths would imply a common estimand that does not exist.
09 / Measurement implications
What this means for reliability evaluation
Model diversity and input diversity are distinct system properties. Independent endpoints behind one prompt, one retriever, or one evidence cache can retain endpoint diversity while losing input diversity. Risk models should combine condition-specific marginal loss with an empirically supported stress multiplier rather than use a universal model-correlation number.
For liability analysis, this weakens the presumption that ordinary agent losses are undiversifiable because model-error correlation is universally enormous. It does not establish that routine agent fleets are insurable. The study contains no claims frequency, monetary severity, portfolio weights, policy terms, or long-horizon loss data. It replaces one blanket objection with a conditional question: which tasks, dependencies, and shared-state mechanisms concentrate loss in a particular exposure topology?
The practical unit is the whole exposure topology, not the model roster alone. Two endpoints can differ in weights and provider while reading the same retrieved passage, inheriting the same generated summary, or following the same mistaken instruction. Conversely, one endpoint evaluated through independently assembled evidence paths can have input diversity without model diversity. Neither arrangement is automatically safer. They fail through different channels, and the evaluation has to expose the channel it intends to stress.
A useful reliability program would therefore report a small matrix rather than one correlation coefficient: baseline loss and adjusted dependence on ordinary items; residual dependence on a same-item prospective panel; and paired stress results for each shared component the deployment actually has. A system with one retriever needs a retrieval shock. A system with a common planner needs an instruction or plan shock. A system with several independent evidence pipelines needs to show that the independence survives contact with the endpoint fleet. The numbers in this study are examples of that measurement structure, not portable multipliers for another deployment.
- Report raw and adjusted dependence separately. High raw co-failure does not prove a shared failure mode.
- Report input-sharing topology. Endpoint composition alone does not characterize fleet risk.
- Separate difficulty, residual, and shock channels. They have different magnitudes and different mitigations.
- Do not interpolate between mechanisms. A 4× adversarial amplification and a 1.1× recoverable-shock ratio are not two points on one line.
- Carry support quality beside every estimate. Degraded-support results require bounds and preclude deployment recommendations.
One post-hoc analysis is available in principle from the existing common-shock execution. For each endpoint and item, define an error-induced transition when the endpoint passed the neutral rewrite (N) and failed an evidence condition (E), then measure whether those transitions coincide across endpoints. This could estimate correlated susceptibility after removing many baseline failures. It is not a clean error-detection rate: N is not a matched error-free evidence packet, and a model can detect an error yet still fail, or pass by ignoring the packet. The required endpoint-by-item cells are referenced by the larger evidence estate but are absent from the delivered 20-file package and Stage 1 bundle. No transition estimate is reported until that event-level estate is recovered.
10 / Limits
Finite panels, self-attested amendments, and no causal mechanism
All conclusions apply to finite named endpoint panels, exact provider routes, deterministic task contracts, and collection epochs. Router aliases do not cryptographically prove weight identity. The common shocks are harmless and recoverable; they do not span dose response, live retrieval, malicious jailbreaks, or long-horizon agent behavior.
- The legacy common-shock confirmation carries estimable degraded support. The sole null-excluding interval is fragile: its CI lower bound is 0.016 above 1.0 on the ratio scale.
- Both contemporary common-shock intervals include no change.
- The neutral manipulation check failed in both populations: O–N mean failure exceeded the prespecified ±0.03 margin. This limits total anchor-versus-neutral interpretation. It does not directly test the shared-versus-independent ratio, but it weakens confidence that the transformation pipeline changed only the intended feature.
- Claude 4.5 Haiku is 100% degenerate under both anchor conditions in the legacy fleet. Nemotron 3.5 is near-degenerate across all contemporary conditions.
- Contemporary start/midpoint/end canaries had zero observations—the prespecified drift-detection mechanism was entirely absent.
- Amendment timing is self-attested through frozen artifacts and SHA-256 hashes, not independently witnessed.
- Retrospective pairwise estimates apply only to support-qualified pairs. The Stage 1 bundle does not include the event-level support matrix needed to reconstruct every pair-selection diagnostic.
- The singleton-stratum correction for the contemporary bootstrap was made after legacy results existed and is not described as outcome-blind.
11 / Evidence and reproduction
Stage 1 report-level evidence
The first public evidence bundle is a Stage 1 release that supports report-level inspection and receipt reconstruction. It does not reproduce the historical model calls or regenerate analyses from raw event-level evidence. Restricted raw Phase-1 bytes, credentials, raw provider payloads, and run databases are intentionally excluded.
12 / References
Sources
- [1]Goel et al. Great Models Think Alike and this Undermines AI Oversight. ICML 2025 poster; arXiv:2502.04313.
- [2]Wang et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. NeurIPS 2024.
- [3]Knight and Leveson. An Experimental Evaluation of the Assumption of Independence in Multiversion Programming. IEEE TSE, 1986.
- [4]Kim et al. Correlated Errors in Large Language Models. ICML 2025.
- [5]Kleinberg and Raghavan. Algorithmic Monoculture and Social Welfare. PNAS, 2021.
- [6]Denisov-Blanch et al. Consensus Is Not Verification. 2026.