When Do Diverse Models Fail Together?
Difficulty, Shared Evidence, and Common Shocks Across Language-Model Fleets
Across a 395-model retrospective panel, median raw agreement fell from 0.498 to −0.018 after item-difficulty adjustment, while empirical failure clusters barely aligned with vendor taxonomy. Evidence-path diversification did not outperform endpoint substitution. Recoverable common shocks produced little stable amplification; severe semantic instruction conflict was a distinct stress mechanism.