CMFC follow-up · Technical note

Methodology and evidence boundaries

What the experiment measured, which record supports it, and where the claims stop.

This note summarizes the working paper and retained evidence. It does not rerun the empirical study, recertify its statistical procedure or replace the paper’s methods and amendment record.

01 / Design and population

Three conditions, one finite route panel

The study uses clean, shared-corruption and IID-corruption conditions over the same planned item families. The two corrupt arms draw from the same pool of complete variants. Shared assignment gives all routes the same variant; IID assignment draws independently with replacement and can therefore contain accidental matches.

A complete variant bundles the error’s false value, position, severity, relevance and salience. The contrast is about assignment of that bundle, not an isolated effect of any one component. The realized IID collision count and topology are absent from the bounded shareable artifacts.

Population and denominator register
QuantityRetained valueScope
Planned item families1,800Arithmetic, constraints, extraction and tool calls; code prospectively excluded
Retained routes8Specific model/provider/deployment endpoints at collection
Primary complete-support families1,677Complete across eight routes and all three conditions
Accepted / planned cells43,073 / 43,20099.71% overall support; not the complete-family denominator
Corrected stage analysis1,716 corrupt; 1,758 cleanSeparate postconfirmatory support sets, not replacements for the primary set

The primary DI contrast uses only shared and IID corruption, although the primary inclusion rule requires clean-arm support too. The exact route identifiers, collection decisions and exclusions remain in the paper. Route names do not independently establish the underlying weights or provider independence.

02 / Outcome definitions

Outputs, not internal mental states

The response contract called for a JSON object containing a decision, target identifier, corrected value, repair object and terminal answer. Non-detection means the corrupted packet was not correctly flagged under that contract. Unsuccessful exact repair means the designated replacement object was not emitted exactly. Terminal task failure means the final answer was incorrect, scored separately.

A correct final answer can coexist with failed exact repair. A valid flag can coexist with a wrong target or replacement. Observable stages therefore do not identify where a model became aware of a problem or where an internal reasoning process failed.

Transport, routing, authentication, ambiguous-delivery and scorer failures remained infrastructure-missing outcomes, not model failures. An accepted provider response violating the behavioral contract could instead be scored as a behavioral failure. Infrastructure retries continued the same logical cell; they were not repeated behavioral draws.

03 / Estimand and inference

Dependence inflation and the primary contrast

For each outcome and corruption condition, DI divides the observed item-to-item variance of the equal-weight route-panel failure fraction by the sum of the route marginal variance contributions under independence. The denominator omits cross-route covariances. Values above one reflect net positive covariance in the tested matrix; values below one can reflect net negative dependence.

The primary contrast is the logarithm of DI under shared assignment divided by DI under IID assignment, reported after exponentiation as a ratio. The saved analysis used 9,999 stratified item-family bootstrap resamples and studentized maximum-absolute-statistic simultaneous intervals across the primary family. Resampling preserved observed counts within strata.

Frozen implemented primary results · Shared / IID DI ratio
OutcomeRatio [simultaneous interval]
Non-detection1.038 [0.948, 1.137]
Unsuccessful exact repair1.129 [1.060, 1.203]
Terminal task failure1.061 [0.992, 1.135]

04 / Protocol and support

Why the omitted gate matters

The support rule required at least 20 projected events and 20 projected non-events for each endpoint, condition and primary outcome. In the retained projection audit, one route had zero projected exact-repair successes in each corrupt arm. Its realized repair successes were three and four. These are sparse outcome counts despite high overall delivery support.

The closeout distinguishes the planned gate from what was actually implemented. The manuscript retains the result of the implemented analysis without treating its null-excluding repair interval as proof that the full protocol succeeded. The closest retained weighting sensitivity remains a postconfirmatory analysis rather than a replacement primary result.

05 / Follow-up results

Different analyses carry different guarantees

Mean unsuccessful repair was 59.23% under IID and 59.91% under shared assignment, a difference of +0.68 percentage points [−0.12, 1.45]. All-eight unsuccessful repair was 77/1,677 (4.59%) versus 117/1,677 (6.98%), a difference of +2.39 percentage points [0.95, 3.82]. These intervals are postconfirmatory, percentile and unadjusted.

The inherited field called “majority failure” uses at least four failures out of eight, not strict majority. Its rate was 69.65% under IID and 68.99% under shared assignment. This is a descriptive threshold comparison, not a new inferential result.

Named equal-mechanism, calibration-composition, condition-specific-support and missing-cell scenarios retained positive repair-dependence contrasts. Their analysis populations and interval procedures must not be pooled with the primary result. The unrestricted binary-completion outer envelope, 0.989–1.311, includes one and is not a confidence interval.

The corrected stage-cascade analysis uses its own support sets and gates. Its terminal estimate cannot upgrade the primary terminal result. The infrastructure analysis uses historical delivery records and is descriptive and noncausal; retry, provider and timing effects are not experimentally separated.

06 / Evidence register

Included reports, referenced raw records

Evidence behind this briefing
LayerIncluded or retainedDoes not establish
D1 follow-up v2Primary report, result tables, sensitivity and stage analyses, infrastructure summaries, analysis code and reviews retained locallyFull raw-data reproduction on this machine
Provenance receiptsManifests, hashes, reproduction and correction recordsScientific validity from file integrity alone
A2-F revised closeoutSynthetic-fixture/API verification: 12 retained tests and 19 fixture casesEmpirical D1 reanalysis or resolution of missing observations
Separate validation handoffBounded completed checks and explicit remaining acceptance workA completed statistical-validation campaign
Restricted evidenceUpstream records referenced by retained manifestsAvailability of prompts, responses, gold state, provider payloads or event databases in this public release

The downloadable PDF is the public working-paper version. Its scientific body is unchanged from the reviewed draft; release notices, availability links and declarations have been finalized. The result downloads contain selected report-level aggregates. They do not replay provider calls or provide a portable analysis environment.

07 / Release posture

Public working paper and declarations

Nicholas Zinner and Beacon Bot, Future Shock. Nicholas Zinner directed the research and authorized publication. Beacon Bot assisted with evidence review, literature retrieval, drafting, revision and document production. Self-funded by Nicholas Zinner; no external funding or competing interests to declare. Human authors remain responsible for scientific interpretation and publication decisions. No external ethics certification is claimed. Restricted upstream records are withheld, and publication does not grant redistribution rights to them.

Available downloads: headline results (CSV), study summary (JSON), evidence notes, and checksums. These files support inspection of reported results, not independent empirical replication.

The scope is a finite synthetic route panel. Neither the paper nor this preview establishes a model ranking, a universal independence coefficient, a deployment-loss multiplier or the prevalence of real-world common shocks.

For the full methods, exact route list, literature and artifact citations, read the paper. For the wider research context, see the original CMFC briefing and its methodology page.

AI news, analysis, and weekly deep dives. No hype.