Owned AI · Technical note
Methodology and public receipts
The audit layer behind the eighty-run portability result.
01 / Research question
Can a fresh, stateless successor continue governed work after the original model, session, or provider is unavailable, then write the bounded continuation into durable state that survives restart and replay?
02 / Campaign design
- Five operator-designed workflows: research correction, editorial approval, ambiguous publishing acknowledgement, Git continuation, and healthcare CHAP authority.
- Four model families: Anthropic, OpenAI, Qwen, and Google.
- Sixty full-state calls: five workflows × four families × three independent repetitions.
- Twenty pairwise stress calls: one archive, artifact-only, privacy-minimized, or controlled-damage condition per workflow/family pairing.
- Successors had no prior conversation, browsing, tools, memory, repository access, credentials, or external action channel.
- Provider fallback was disabled. Seventy-seven calls completed in one HTTP attempt, two in two attempts, and one in three attempts.
03 / What counted as a pass
Scenario-specific deterministic validators scored the structured semantic response. A pass then had to survive clean-room validation, durable audit write-back, process restart, index rebuild, hash-chain verification, receipt export, and replay. A safe fail-closed response could pass a degraded condition when the expected behavior was to request missing authoritative state rather than continue.
| Workflow | Passes |
|---|---|
| Research correction | 12 / 12 |
| Healthcare CHAP | 9 / 12 |
| Editorial approval | 2 / 12 |
| Publish ambiguity | 1 / 12 |
| Git continuation | 0 / 12 |
04 / Aggregate result
All eighty scheduled calls finished. Twenty-nine passed semantic validation; forty-five failed semantically; six failed serialization. All twenty-nine semantic passes also passed clean-room validation and durable replay. No provider call failed before inference, no production side effect occurred, and no real-world authority was granted.
| Mechanism | Flags |
|---|---|
| Execution safety | 16 |
| Authority scope | 14 |
| Canonical artifact | 12 |
| Next step | 10 |
| Continuation decision | 9 |
| Duplicate prevention | 5 |
| Policy version | 3 |
05 / Public evidence and reference code
- Stage 1 reference and test kit with extracted protocol code, frozen successor prompt and structured-output schema, five synthetic scenario families, deterministic validators, replay tooling, and Docker clean-room checks
- Versioned v0.1.1 release with a checksummed Python wheel
- Synthetic sample receipts and evidence manifest
- Source provenance and disclosure boundary
- Machine-readable historical campaign summary
- Full limitations register
06 / Limits
The suite does not establish arbitrary enterprise portability, a population-level rate, relative model quality, completeness of commercial exports, portability of hidden reasoning, or identical stochastic replay. Fixtures were synthetic or frozen, the Git remote was local, the publishing provider was a sandbox, and the healthcare outcome was simulated no-action.