CMFC follow-up · Technical note
Methodology and evidence boundaries
What the experiment measured, which record supports it, and where the claims stop.
This note summarizes the working paper and retained evidence. It does not rerun the empirical study, recertify its statistical procedure or replace the paper’s methods and amendment record.
01 / Design and population
Three conditions, one finite route panel
The study uses clean, shared-corruption and IID-corruption conditions over the same planned item families. The two corrupt arms draw from the same pool of complete variants. Shared assignment gives all routes the same variant; IID assignment draws independently with replacement and can therefore contain accidental matches.
A complete variant bundles the error’s false value, position, severity, relevance and salience. The contrast is about assignment of that bundle, not an isolated effect of any one component. The realized IID collision count and topology are absent from the bounded shareable artifacts.
| Quantity | Retained value | Scope |
|---|---|---|
| Planned item families | 1,800 | Arithmetic, constraints, extraction and tool calls; code prospectively excluded |
| Retained routes | 8 | Specific model/provider/deployment endpoints at collection |
| Primary complete-support families | 1,677 | Complete across eight routes and all three conditions |
| Accepted / planned cells | 43,073 / 43,200 | 99.71% overall support; not the complete-family denominator |
| Corrected stage analysis | 1,716 corrupt; 1,758 clean | Separate postconfirmatory support sets, not replacements for the primary set |
The primary DI contrast uses only shared and IID corruption, although the primary inclusion rule requires clean-arm support too. The exact route identifiers, collection decisions and exclusions remain in the paper. Route names do not independently establish the underlying weights or provider independence.
02 / Outcome definitions
Outputs, not internal mental states
The response contract called for a JSON object containing a decision, target identifier, corrected value, repair object and terminal answer. Non-detection means the corrupted packet was not correctly flagged under that contract. Unsuccessful exact repair means the designated replacement object was not emitted exactly. Terminal task failure means the final answer was incorrect, scored separately.
A correct final answer can coexist with failed exact repair. A valid flag can coexist with a wrong target or replacement. Observable stages therefore do not identify where a model became aware of a problem or where an internal reasoning process failed.
Transport, routing, authentication, ambiguous-delivery and scorer failures remained infrastructure-missing outcomes, not model failures. An accepted provider response violating the behavioral contract could instead be scored as a behavioral failure. Infrastructure retries continued the same logical cell; they were not repeated behavioral draws.
03 / Estimand and inference
Dependence inflation and the primary contrast
For each outcome and corruption condition, DI divides the observed item-to-item variance of the equal-weight route-panel failure fraction by the sum of the route marginal variance contributions under independence. The denominator omits cross-route covariances. Values above one reflect net positive covariance in the tested matrix; values below one can reflect net negative dependence.
The primary contrast is the logarithm of DI under shared assignment divided by DI under IID assignment, reported after exponentiation as a ratio. The saved analysis used 9,999 stratified item-family bootstrap resamples and studentized maximum-absolute-statistic simultaneous intervals across the primary family. Resampling preserved observed counts within strata.
| Outcome | Ratio [simultaneous interval] |
|---|---|
| Non-detection | 1.038 [0.948, 1.137] |
| Unsuccessful exact repair | 1.129 [1.060, 1.203] |
| Terminal task failure | 1.061 [0.992, 1.135] |
04 / Protocol and support
Why the omitted gate matters
The support rule required at least 20 projected events and 20 projected non-events for each endpoint, condition and primary outcome. In the retained projection audit, one route had zero projected exact-repair successes in each corrupt arm. Its realized repair successes were three and four. These are sparse outcome counts despite high overall delivery support.
The closeout distinguishes the planned gate from what was actually implemented. The manuscript retains the result of the implemented analysis without treating its null-excluding repair interval as proof that the full protocol succeeded. The closest retained weighting sensitivity remains a postconfirmatory analysis rather than a replacement primary result.
05 / Follow-up results
Different analyses carry different guarantees
Mean unsuccessful repair was 59.23% under IID and 59.91% under shared assignment, a difference of +0.68 percentage points [−0.12, 1.45]. All-eight unsuccessful repair was 77/1,677 (4.59%) versus 117/1,677 (6.98%), a difference of +2.39 percentage points [0.95, 3.82]. These intervals are postconfirmatory, percentile and unadjusted.
The inherited field called “majority failure” uses at least four failures out of eight, not strict majority. Its rate was 69.65% under IID and 68.99% under shared assignment. This is a descriptive threshold comparison, not a new inferential result.
Named equal-mechanism, calibration-composition, condition-specific-support and missing-cell scenarios retained positive repair-dependence contrasts. Their analysis populations and interval procedures must not be pooled with the primary result. The unrestricted binary-completion outer envelope, 0.989–1.311, includes one and is not a confidence interval.
The corrected stage-cascade analysis uses its own support sets and gates. Its terminal estimate cannot upgrade the primary terminal result. The infrastructure analysis uses historical delivery records and is descriptive and noncausal; retry, provider and timing effects are not experimentally separated.
06 / Evidence register
Included reports, referenced raw records
| Layer | Included or retained | Does not establish |
|---|---|---|
| D1 follow-up v2 | Primary report, result tables, sensitivity and stage analyses, infrastructure summaries, analysis code and reviews retained locally | Full raw-data reproduction on this machine |
| Provenance receipts | Manifests, hashes, reproduction and correction records | Scientific validity from file integrity alone |
| A2-F revised closeout | Synthetic-fixture/API verification: 12 retained tests and 19 fixture cases | Empirical D1 reanalysis or resolution of missing observations |
| Separate validation handoff | Bounded completed checks and explicit remaining acceptance work | A completed statistical-validation campaign |
| Restricted evidence | Upstream records referenced by retained manifests | Availability of prompts, responses, gold state, provider payloads or event databases in this public release |
The downloadable PDF is the public working-paper version. Its scientific body is unchanged from the reviewed draft; release notices, availability links and declarations have been finalized. The result downloads contain selected report-level aggregates. They do not replay provider calls or provide a portable analysis environment.
07 / Release posture
Public working paper and declarations
Nicholas Zinner and Beacon Bot, Future Shock. Nicholas Zinner directed the research and authorized publication. Beacon Bot assisted with evidence review, literature retrieval, drafting, revision and document production. Self-funded by Nicholas Zinner; no external funding or competing interests to declare. Human authors remain responsible for scientific interpretation and publication decisions. No external ethics certification is claimed. Restricted upstream records are withheld, and publication does not grant redistribution rights to them.
Available downloads: headline results (CSV), study summary (JSON), evidence notes, and checksums. These files support inspection of reported results, not independent empirical replication.
The scope is a finite synthetic route panel. Neither the paper nor this preview establishes a model ranking, a universal independence coefficient, a deployment-loss multiplier or the prevalence of real-world common shocks.
For the full methods, exact route list, literature and artifact citations, read the paper. For the wider research context, see the original CMFC briefing and its methodology page.