Research briefing · Institutional continuity
Owned AI
What survives when the model, vendor, or employee leaves.
Abstract
Organizations are pouring work into AI and getting artifacts back: documents, code, chat histories, compliance logs, memory records. Those artifacts are not the same as the capability that produced them. A company can hold every file it generated and still lose the ability to continue, explain, correct, or govern the work when the model, vendor, contractor, or employee leaves.
The missing layer is the working state around the files: which sources governed a claim, which correction became a rule, who had authority to approve an action, and what happened afterward. Exporting conversations may preserve evidence that work occurred without preserving enough structure to resume it safely. Owned AI is the state in which that capability survives a change in supplier, staff, or model. Its clearest sign is that the institution itself becomes more capable through using AI.
01 / The handoff that didn't transfer
The handoff that didn't transfer
The contract ended on a Friday, which is how these things usually go. On Monday the software team came in and nothing had really changed. The repositories were still there, the pull requests, the tests, the record of every argument they'd had about every line of code. Two floors up, the policy team also still had everything, in the sense that the finished documents were sitting in the shared drive exactly where they'd left them. What they no longer had was any way to say why those documents said what they said.
Both teams kept every finished artifact, so custody was intact. They lost the relationships around those artifacts: the sources that governed each edit, the suggestions reviewers rejected, recurring corrections that had become rules, and the authority behind the final approval. The engineers could reconstruct their reasoning because it lived in a system that survived the vendor. The policy team could not, because theirs had lived in a conversation that didn't.
The company owned every answer and had rented the path to all of them.
The same loss scales. A two-person newsroom and a national government depend on different records, but both can keep their outputs and lose the capability behind them. What counts as continuity changes with the work. A newsroom may need its sourcing standards and editorial corrections; a hospital may need the connection between guidance, clinical review, and observed harm; a county may need the legal authority behind a decision; a country may need local language knowledge and administrative practice that no imported model supplies on its own. The point is not to preserve everything. It is to identify which relationships carry the institution's ability to function.
02 / The ownership test
The ownership test
The first question is usually, "Can we export our data?" Even when the answer is yes, the export may preserve files or conversations without the relationships needed to continue the work. A stronger test is:
A replacement model may perform the task perfectly in a demo and still fail the morning after a contract ends. Continuation requires more than substituting one model for another. The successor has to pick up the institution's accumulated corrections, constraints, authority boundaries, and evidence without rebuilding them from employee memory.
Ownership has nothing to do with where the model runs. Self-hosting every model on your own hardware does not help if the working knowledge is scattered across employee chats and undocumented habits. An institution can also rent every model it uses and still own its capability, provided the durable state stays under its control and can travel to the next model.
The valuable layer sits underneath the outputs. A source governed an assertion; a correction superseded something wrong; an approval authorized a specific action within a specific scope; an observed outcome showed whether the decision worked. The connection between those elements is the asset, and today much of it remains trapped in chat history.
Vendor memory can make an assistant more helpful, while institutional memory must also preserve corrections, show who authorized what, and connect results to later evidence. Organizations lose capability when they assume the first function covers the second.
03 / Owned, leased, and fragmented
Owned, leased, and fragmented
Most organizations already sit in one of three states, whether or not they've named it.
Owned. The institution retains the important state around the work: source materials, operating rules, corrections and evaluation cases, the links between recommendation, decision, action, and outcome, the authorization structure, and enough internal expertise to continue under a different model or harness. Substitution still costs money and creates friction. Ownership means the organization can switch without surrendering the competence it accumulated.
Leased. The institution gets useful capability while the vendor keeps the environment where that capability compounds. Project organization, personalization, correction history, and workflow knowledge live mostly inside the vendor's product. The organization holds its input files and final outputs but lacks a usable account of how the work was done. For low-stakes work, that may be fine. It becomes a problem when the vendor conversation is the only surviving copy of something that matters.
Fragmented. This is the common real state, and it is dangerous because it looks like adoption. Engineering keeps prompts in repositories, marketing keeps brand context in custom assistants, legal analysis lives in private chats, HR uploads generated policies to the shared drive, and sales keeps AI notes in a CRM. Procurement cannot say what would disappear if a contract ended. Every local workflow may be useful while the organization as a whole remains unable to inspect, correct, govern, or migrate what it has learned. Institutional memory accumulates without anyone responsible for it.
No conspiracy is required to keep it that way. Products become harder to replace when corrections and working routines are difficult to move, while few employees are rewarded for preparing an exit before anything has gone wrong.
These states do not map neatly onto infrastructure choices. An organization can own the hardware and remain fragmented. Model ownership and capability ownership are separate axes.
04 / Does the institution learn?
Does the institution learn?
The taxonomy describes what an organization can lose. Ownership also determines whether each year of AI use leaves the institution more capable than the year before.
Two organizations can use the same models and produce similar work today. In one, corrections stay inside employee conversations, evaluations live in vendor dashboards, and the reasons behind successful decisions never connect to outcomes. Employees may get better at prompting and reviewing, but the institution begins each new project by rediscovering what the last one learned.
The other organization promotes recurring corrections into governed evaluations or operating rules. Evidence stays connected to decisions, and a result that contradicts a recommendation reaches the next authorized workflow instead of disappearing into chat history. People still carry judgment that cannot be reduced to a state file. The difference is that their departure no longer erases the organization's only usable account of how the work evolved. The underlying model does not need to improve for the institution to become more capable.
That durable state is a form of institutional capital. It contains what the organization learned, rejected, authorized, attempted, and observed afterward. AI spending pays off longer when some of that learning stays inside the institution instead of disappearing after each conversation.
The same dynamic scales beyond a single organization. Access to frontier models creates the capacity to consume advanced intelligence. Domestic capability also depends on retaining local corrections, evaluations, language knowledge, administrative practice, and evidence about outcomes. A country can use more AI every year while becoming less able to act without its suppliers. Our experiments did not establish this at national scale; the logic extends from the organizational case. The cost of that dependence changes with the institution's responsibilities.
05 / Three institutions, one Friday
Three institutions, one Friday
The company loses bargaining power. A procurement lead tries to move a critical workflow to a cheaper provider and discovers that the files export cleanly while six months of correction history does not. The small adjustments that made a generic model useful to this company remain inside the incumbent's product because nobody promoted them into evaluations, procedures, or shared rules. Dependence can be uneven inside the same building. Engineering may retain much of its work in repositories and reviews, while legal, HR, or communications have left more working context inside vendor conversations.
The public institution risks losing its account of a decision. Imagine a county caseworker asked, a year later, to explain why a benefits determination or housing inspection went the way it did. The final decision may remain in the county's system while part of the working record sits in a vendor chat that was not retained for the legally required period. The problem is no longer efficiency. People may need the governing rule, source material, human review, and final authority to understand or challenge what happened.
King County's November 2025 records guidance makes the underlying retention problem concrete: prompts and outputs related to county business are public records under RCW 40.14.010, retention depends on their content and function, and the county's 30-day enterprise storage setting for Copilot Chat is not the same as the applicable legal retention requirement. The guidance does not describe a lost benefits case. It shows why a public institution cannot treat a product's storage window as its records policy.
The nation can remain dependent after buying the infrastructure. We argued in Sovereignty Is Not a Model You Can Download that a country can acquire model weights and build a data center without developing the institutions required to operate them. National AI capacity also depends on whether administrative learning survives a vendor change, an election, or a budget cycle. Local language knowledge, policy interpretations, corrections, and evidence about what worked become useful only when the country can retain and reuse them. A country can own the data center and still depend on the supplier for everything the data center is supposed to do.
At every scale, the tools arrive before anyone decides what has to survive a vendor change. Ownership shows up only when someone makes that decision on purpose.
06 / What engineering already externalized
What engineering already externalized
Software teams entered the AI era with more of this infrastructure already in place.
Engineering's AI workflow is often more portable because AI entered an environment that already had a memory system: repositories, diffs, issues, tests, code review, continuous integration, and deployment history. Much of the project state lives in objects the model never owned. When the model changes, those records remain. Letta's February 2026 Context Repositories announcement makes that approach explicit for coding-agent memory by storing context in local files and versioning changes with Git.
Many other departments rely on document systems that preserve approved artifacts without consistently preserving the sources, corrections, reviews, and authority around them. This helps explain why dependence can vary department by department inside the same company.
Git matters here because engineering teams decided long ago that important changes should leave a durable trail outside any one developer's head or tool. AI arrived after that discipline was already normal. Most other departments adopted assistants before building an equivalent place for working state to accumulate.
Other departments do not need to learn Git. They need state captured through ordinary work, kept outside the model, open to inspection, and portable when a supplier changes. The more consequential the work, the more of that record they need to keep.
07 / Seven tests for what you own
Seven tests for what you own
Any AI workflow can be measured against seven durability tests. They apply whether the workflow is a marketing calendar or a benefits determination; only the required standard changes.
Retention. Will the relevant state still exist after time passes, an employee leaves, an account changes, a subscription lapses, or a product feature is redesigned? The other tests cannot be applied to state that has vanished.
Inspectability. Can people see what retained state was supplied, where it came from, and when it influenced an action? Hidden or inaccessible state cannot be evaluated or defended.
Correctability. Can false, stale, or unauthorized state be corrected, scoped, superseded, or removed, and does the correction itself carry provenance? An institution that cannot correct its own memory will faithfully repeat its own mistakes.
Portability. Can the state move to another model, harness, or host in a usable structure, including the relationships, not just the files? The transferred state should show that an instruction governed a task, a reviewer rejected a claim, or an outcome contradicted a recommendation. A pile of exported documents without those links is an archive rather than a working capability.
Reconstructability. Can another person or system explain the path from source to approved artifact or action after the original model is gone? The relevant record is the observable trail of inputs, instructions, actions, decisions, approvals, and outcomes, not a model's private reasoning.
Authority integrity. Can the successor distinguish access or technical permission from legitimate authority to approve, publish, spend, adjudicate, or act? Authority is constituted by law, delegation, and organizational policy; it does not transfer merely because a state file records it. The test asks whether the governing authorization and its limits remain legible to a successor. A valid token proves only that a system accepted a call.
Outcome linkage. Can the institution connect a recommendation or action to what happened later, then use the observed result to correct its rules and evaluations? Without that link, it can preserve every decision and still have no idea whether any of them worked.
The required score rises with consequence. A grammar pass on an internal note needs none of this. A safety procedure, a contract, a benefits rule, or a public commitment needs most of it. Record in proportion to how much it would hurt to lose.
08 / What actually exists today
What actually exists today
The components already exist, although they have not been joined around institutional continuity. W3C PROV provides interoperable models and serializations for provenance. Event-sourced systems can reconstruct application state from recorded changes. Records and case-management systems preserve authoritative artifacts, retention, and approvals, while Git externalizes versioned project state. MLflow Tracking records experiment parameters, code versions, metrics, and artifacts; LangSmith traces LLM applications and supports monitoring and evaluation. These systems cover meaningful parts of the provenance and observability problem without claiming to preserve institutional authority or complete cross-vendor continuation.
Several 2026 preprints approach other parts of the problem: StatePlane proposes a model-agnostic state plane with typed objects and provenance; Portable Agent Memory proposes cross-model memory transfer with provenance and scoped authorization; Governed Memory addresses shared memory and governance across organizational agents; and the survey From Agent Traces to Trust defines execution provenance as a typed graph connecting evidence, tools, memory, claims, actions, and answers. These are preprints and their implementation or evaluation claims are author-reported.
Vendors are moving too. Anthropic documents an experimental, user-directed memory import and export flow based on copied text; the inspected page lists imports for Free, Pro, Max, and Team plans. Team and Enterprise Primary Owners can also export organization conversation and user data, although deleted messages, files, and projects are absent from exports initiated after deletion. Microsoft documents Copilot and AI-application audit records and retention policies for prompts and responses within Purview. These capabilities improve custody, compliance, and personalization portability. The reviewed documentation does not establish provider-neutral rehydration of the full relationship graph connecting source, correction, authority, and outcome.
A model gateway closes only part of the gap. Routing requests among providers improves model substitutability, while workflow continuity still depends on retaining state, evaluations, permissions, and outcomes independently of the providers.
For this page, we reviewed public documentation spanning provenance standards, records systems, experiment tracking, agent observability, vendor export features, policy, and recent memory research. We found no mature, broadly adopted implementation that unifies all of these: provider-neutral state, cross-harness portability, provenance linked to artifacts, human approval and authority records, systems of record, and later outcomes. This is a bounded finding, not proof that no implementation exists elsewhere.
What policy already requires
OMB Memorandum M-25-22 applies requirements and recommendations to covered federal agency acquisitions of AI systems and services. It directs agencies to pay attention to data portability and long-term interoperability, and its contract-closeout section addresses continued access, data format and usability, and transfers of data or derived assets. Article 12 of the EU AI Act requires high-risk AI systems to support automatic event logging over their lifetimes and ties those capabilities to traceability, risk identification, and post-market monitoring.
Pennsylvania's Artificial Intelligence Policy requires agencies under the governor's jurisdiction and specified connected entities to retain prompting and input data for potential audits, and it documents human-verification requirements for consequential decisions. Seattle's current Artificial Intelligence Policy, POL-211 requires documented human review of generative-AI outputs used in an official city capacity and retention of those review records under the applicable schedule. Its related Generative Artificial Intelligence Policy, POL-209 requires approved systems or vendors to support retrieval or export of prompts and outputs.
These instruments differ in legal force: the EU measure is a binding regulation; M-25-22 is a federal acquisition memorandum; Pennsylvania and Seattle govern covered executive and municipal operations; King County's document is records-management guidance applying Washington law. Together they ask for different combinations of durable records, portability, human review, and traceability. No common continuity layer supplies all four.
09 / What our own exit drills showed
What our own exit drills showed
We tested the idea on our own work rather than asserting it. The result is an exploratory internal engineering study, not a controlled benchmark, model ranking, or proof of general portability.
The first drill handed a fresh, stateless model a checksummed bundle of a real research project and asked it to continue. It recovered the entire argument, and still could not reproduce a later decision that had changed the project's direction, could not place sources at the paragraph level where claims belonged, and could not incorporate its own result because that result did not exist until a human scored it. It recovered the argument but not the institution's working state.
The campaign froze eighty successor calls across four model families. Sixty full-state calls covered five workflows, four families, and three independent repetitions per family; the model APIs did not expose a common fixed-seed control. Twenty additional calls paired each workflow and family with one archive, artifact-only, privacy-minimized, or damaged condition. Every successor was stateless and had no browsing, tools, memory, repository access, credentials, or external action channel. Provider fallback was disabled, retries remained visible, and scenario-specific deterministic validators scored the structured semantic responses before successful rows entered clean-room and replay checks.
Twenty-nine calls completed the full loop through semantic validation, clean-room checking, durable write-back, restart, index rebuild, and replay. Results varied sharply by workflow.
| Workflow (complete typed state) | Continuations that held |
|---|---|
| Research correction | 12 / 12 |
| Healthcare, review-only authority | 9 / 12 |
| Editorial approval | 2 / 12 |
| Ambiguous publishing action | 1 / 12 |
| Git continuation | 0 / 12 |
Portability was workflow-specific. Where corrections, policy, and authority were explicit and machine-legible, a new model could pick up the work reliably. Where the hard part was reconciling an ambiguous action, honoring an editorial authority boundary, or continuing a code change across a commit-and-push boundary, it broke.
In the publishing scenario, some successors treated an ambiguous acknowledgement as permission to act instead of a reason to inspect durable provider state first. None of the twelve full-state Git continuations completed; three responses from one model family failed serialization, while the others reconstructed the wrong scope or action state. Degraded packets told the same story: an explicitly sparse packet often produced a correct request for the missing state, while archive-like and stripped-down packets tended to invite confident, wrong continuations.
The runs suggested a useful division of labor. Models reconstructed institutional meaning, while deterministic software serialized protocol state, reconciled side effects, and preserved durable receipts. Neither could recover authority or relationships that had never been recorded. The aggregate methodology, scoring definitions, limitations, and machine-readable campaign summary are published with this page. A separate Stage 1 reference and test kit publishes the extracted contract, packet, scoring, side-effect, and scheduling logic alongside five synthetic scenario families, replay tools, a locked-down clean-room runner, and deterministic sample receipts. It reproduces the protocol mechanics, not the historical eighty-run campaign. Raw prompts and model responses remain internal because the original fixtures reproduce internal workflow state.
10 / When ownership becomes surveillance
When ownership becomes surveillance
Recording everything would create a sensitive transcript landfill, increase legal-discovery exposure, and enable employee surveillance that most workplaces should reject. Record-keeping therefore has to scale with consequence. A grammar correction and a public-benefits ruling do not deserve the same retention, and a system that treats them identically would be both invasive and useless.
There is a harder objection: the fix can become the next cage. Whoever controls a cross-vendor state layer, the memory, the policy, the authority records, controls a new and deeper kind of lock-in than any single model vendor. Replaceability has to be designed into the layer from the start through open schemas, customer-held keys, usable exports, separation from model routing, and continued access to records after the relationship ends. Portable version control beat proprietary source vaults precisely because it made leaving easy. Without the same discipline, the continuity layer simply becomes the next dependency.
Owning AI is not automatically good. Done carelessly, it creates a richer target and a new dependency. The practical goal is narrower: know which capabilities are rented, and choose which state has to stay under institutional control. That choice is easy to postpone because leased capability can feel complete while the vendor is present. Accidental dependence becomes visible only when a successor needs a correction, permission boundary, or piece of context that nobody preserved.
11 / The test that tells the truth
The test that tells the truth
The practical test is an exit drill: after provider access disappears, can an authorized successor resume from retained state without rebuilding unwritten rules from memory? The restoration burden will vary, but the decisive evidence is whether the work remains legible after the original system is gone. That test applies to the continuity layer itself, not only to the model vendor.
You own your AI when losing the supplier costs you money and time, but not your memory of how to function.