If your organisation changed AI provider today, which parts of its work would carry over? Documents may be easy to move. Accumulated context, working integrations and reliable behaviour may be harder to reproduce.
This briefing uses sovereign intelligence to mean the ability to choose and replace AI systems while retaining useful company knowledge and control over access. That ability depends on where information lives, how it is accessed and whether another system can use it effectively.
The sections below review research on cost and capability, explain the parts of an agent and offer a checklist for assessing switching costs. The research findings and the proposed assessment framework are separate; this is not a new empirical study.
The economics keep changing
Epoch AI estimates that the cost of reaching a given benchmark performance fell roughly 47% per quarter over 2023–26. Its historical comparison puts that pace ahead of DNA sequencing, compute, lithium batteries and electricity. These are different technologies measured over different periods, not competing products observed in a controlled experiment.
For a CIO, the implication is architectural: a fixed choice of model can become economically outdated while the surrounding implementation is still being rolled out. The relevant question is how much surrounding work must change when the model changes.
The pace of falling prices
Relative cost on a logarithmic scale. Historical average trends reconstructed from the source rates, with each period starting at 1. AI uses the 2023–26 estimate; this chart does not extend any series beyond its study period.
Epoch AI · Figure 1 ↗Method and underlying data
For each technology, relative cost = (1 − quarterly decline)4 × elapsed years. Lines show compounded average rates, not individual price observations. The plot clips at a 100,000-fold reduction. Different periods and units limit direct comparisons.
| Technology | Period | Quarterly decline |
|---|---|---|
| LLM inference | 2023–26 | 47.0% |
| DNA sequencing | 2001–22 | 14.2% |
| Compute | 1940–2001 | 9.9% |
| Lithium batteries | 1991–2024 | 3.6% |
| Electricity | 1892–1973 | 1.2% |
Lower inference prices do not automatically lower the total cost of a workflow. Tool usage, retries, integration and human review also count. Measure cost per successfully completed task, alongside quality, before changing the system.
The frontier is moving in two directions
The cost of a particular capability can fall while the best available capability improves. These are separate developments. A cheaper substitute for yesterday’s model is not the same thing as a system that can solve a harder problem today.
The Epoch Capabilities Index brings benchmark results onto a common scale. The snapshot below shows both individual models and the highest score available over time. It provides context for periodic evaluation, but does not establish which model is suitable for a particular business workflow.
A rising benchmark frontier
Epoch Capabilities Index
Snapshot: 28 September 2026
Points are model estimates; the line is their running maximum. ECI is a benchmark index, not business productivity. Confidence intervals are omitted for readability and are available in the downloadable snapshot.
Epoch AI · ECI ↗Download data ↓A benchmark score is not a percentage of business tasks completed or a promise of productivity. For routine work, a smaller model may be sufficient. For difficult work, assess what the strongest eligible models can do before designing a process around a weaker model’s limitations.
Four terms, four different decisions
Conversations about ‘the AI’ often bundle several decisions together. Separating the layers makes it easier to ask what you are buying, what you are configuring and what you need to control.
The agent is the combined model and harness at work, not an extra piece of software wrapped around them. The environment is what that system can access: documents, applications, databases and other services. Our diagram uses nested boundaries to explain these relationships, not to prescribe a deployment topology.
In an integrated frontier-lab platform, the provider controls both the model and the harness that runs it. You configure and use those layers; you do not necessarily own or control them. That arrangement is a product choice, not a requirement of the architecture: a separate harness can connect to models from different providers.
From capability to a working system
The world the agent can observe and change.
Files, knowledge, applications and services reached through tools. The harness runs the interaction; the environment is what it interacts with.
Select a layer to explore its role. The agent contains the model and harness as a combined system.
Terminology adapted from the AI Coding Dictionary ↗Access also has a time dimension. Reading a file gives an agent a snapshot. If the file changes, the agent needs to read it again. Connecting a system is only the beginning; useful environments also make current information discoverable and clearly distinguish it from stale material.
Better know-how can improve the same model
The WikiSkill preprint studies how agent experience can become persistent knowledge and reusable skills. In its Gemini-3.5-Flash comparison, average accuracy across five benchmarks rises from 49.5% without skills to 68.1% with WikiSkill: an increase of 18.6 percentage points, with the same inference model.
That is evidence for the value of structured, reusable know-how. It is not a measured uplift from adding proprietary company data, and it should not be presented as a forecast of enterprise performance. The reported results average three independent runs of the skill-evolution process.
The same model, with better know-how
Gemini-3.5-Flash · average accuracy across five benchmarks. An experimental skill-evolution result, not a measured gain from proprietary company data.
WikiSkill · Table 1 ↗The business hypothesis is worth testing: your definitions, worked examples and validated procedures may make an agent more effective on your tasks. A comparison on held-out work can test that hypothesis. More documents do not necessarily produce more accurate results.
Build the environment the company keeps
Owning the model and harness gives a frontier lab two reasons to keep work inside its platform: more model usage and a deeper role in the customer’s operations. Memory, procedures and connections can make the product more useful while also increasing the work required to leave. This creates a commercial incentive for company knowledge to accumulate inside the platform.
The alternative is to make the company’s environment the permanent home for the services agents rely on: MCP tools, long-term memory, knowledge bases, reusable procedures and messaging between agents. In this design, the model and harness access those services; they are not the only place those services or their records exist.
The five areas below turn that principle into enterprise responsibilities. These services can be self-hosted or managed: what matters is control over access, usable records and the ability to change providers. The suggested owners should fit your organisation’s existing accountabilities.
Your environment stays with you
Replaceable agent
Model + harness of your choice
A shared company environment
- MCP toolsReusable access to business systems.
- Long-term memoryDecisions and context retained across sessions.
- Knowledge basesSource documents, definitions and references.
- Skills and proceduresVersioned instructions and ways of working.
- Agent messagingShared inboxes, task records and handoffs.
- Permissions and auditAccess policies and durable action records.
Agents change. Shared knowledge, tools and coordination remain.
A proposed architecture: keep durable state and shared services outside a particular harness. Compatible agents connect to the same environment. MCP does not itself provide agent messaging or guarantee portability.
Knowledge and memory
- What belongs here
- Knowledge bases, source documents, customer and case history, agreed definitions, and durable records of decisions and preferences. Keep provenance and retention rules with the information.
- Why it matters for independence
- A replacement agent can recover the organisation’s context without people having to teach it everything again. Its immediate working context still lives in the harness; durable memory needs an explicit read and write mechanism.
Suggested ownerBusiness knowledge owners and the data team
Tools and business systems
- What belongs here
- Company-controlled APIs and MCP servers that expose approved actions in CRM, ERP, document stores and other operational systems. Credentials and access rules belong with these services.
- Why it matters for independence
- Different compatible harnesses can use the same business capabilities. You adapt the agent’s connection instead of recreating every integration. The MCP client remains in the harness; the server exposes the company’s tools and information.
Suggested ownerEnterprise IT and integration teams
Skills and ways of working
- What belongs here
- Versioned procedures, reusable skills, instructions, worked examples and escalation rules. Store the authoritative versions in repositories or document systems the company controls.
- Why it matters for independence
- Your operational know-how survives a change of AI product. A new harness may need different packaging or prompts, but the underlying procedure and its history remain available to reuse and test.
Suggested ownerProcess owners and domain specialists
Messaging and coordination between agents
- What belongs here
- Shared inboxes, message history, task queues, assignments, handoff records and persistent workflow status. Record who owns each task and what has already happened.
- Why it matters for independence
- Agents from different providers can coordinate through shared services. A replacement can pick up a recorded handoff instead of losing the work inside a platform’s private conversation. Active sessions do not transfer automatically; messaging requires its own service, not just MCP.
Suggested ownerWorkflow owners and the automation team
Permissions, audit and evaluation
- What belongs here
- Identity and access policies, approval requirements, durable action logs, evaluation cases and acceptance criteria. Enforce access at the services agents use, as well as in the harness.
- Why it matters for independence
- You retain the evidence needed to compare a replacement and reconstruct past actions. Verify that the new harness enforces approvals and produces the required records before letting it take over.
Suggested ownerSecurity, risk and the accountable business owner
Consider replacing the agent handling a customer request. The replacement can retrieve the same case history, use the same business-system tools and pick up a recorded handoff from another agent. The company retains the shared records and services; the replacement still needs compatible connections, permissions and an evaluation on the work. Active sessions and model behaviour do not automatically transfer. That is the practical independence this architecture aims for: changing the intelligence without recreating the company’s operating context.
What would you lose if you switched provider?
Choose one workflow and name both the current system and a possible replacement. Specify whether the change covers just the model, the harness, or the entire platform. Switching a model within an existing harness is a different exercise from moving the whole workflow.
For each item below, record what would happen in that specific move: it carries over, needs rebuilding, is lost, or remains unknown. Assign an owner and estimate the effort and disruption. An export option is useful evidence only after the exported material has been used successfully in the replacement.
CIO switching checklist
Assess the dependency, then ask for evidence
“Carries over” means usable in the replacement. “Needs rebuilding” means recoverable with work. “Lost” means unavailable in the proposed move. Leave untested assumptions as “Unknown”.
- 01
Company records and source material
Can you take the original documents, metadata and source references with you?
Potential loss. Missing material, provenance or usable access rules.
Evidence to request
Import a representative export into the replacement and check completeness, links and permissions.
- 02
Accumulated context and memory
What has the system learned about your work that exists only in its memory or conversation history?
Potential loss. Continuity: people may have to explain the same work again.
Evidence to request
Recover decisions, preferences and relevant history, then check whether the replacement can retrieve and use them.
- 03
Instructions, skills and workflows
Which procedures can you reuse, and which depend on this harness?
Potential loss. Working routines and the effort already spent refining them.
Evidence to request
Run a documented procedure in the alternative. Record changes to prompts, tools and approval steps.
- 04
Business-system connections
Which integrations and scheduled jobs would stop working?
Potential loss. Access to operational systems or uninterrupted automation.
Evidence to request
Exercise required read and write operations, authentication and error handling in a test environment.
- 05
Permissions, approvals and audit history
Can you retain the controls and evidence required for this workflow?
Potential loss. Control coverage or the ability to reconstruct past actions.
Evidence to request
Check role mappings, approval paths and accessible historical logs. Verify the replacement enforces the same boundaries.
- 06
Evaluations and feedback
Do you own the examples and feedback that establish whether the system works?
Potential loss. A reliable quality baseline and accumulated feedback.
Evidence to request
Use the same held-out cases and acceptance criteria in both systems. Keep results outside either platform.
- 07
Model-specific behaviour
Which results depend on this model, a fine-tune or a proprietary feature?
Potential loss. Task performance, even when the same documents and instructions transfer.
Evidence to request
Compare difficult cases, output formats and failure modes. Verify whether any custom model assets are actually transferable.
- 08
Day-to-day operation
How much interruption, retraining and reconfiguration would the move create?
Potential loss. Time, service continuity and the team’s familiarity with the system.
Evidence to request
Run a limited parallel trial. Record staff time, review load, service dependencies and a rollback plan.
Selections stay in this tab and reset on reload. The Wavyr PDF includes your selections and space for evidence, an owner and estimated effort. It is created in your browser. These counts are an inventory, not a risk score.
Treat unknowns as unanswered questions, not evidence of portability. Before deciding to switch, run the same held-out work in both systems and compare result quality, review effort and total cost. A dependency may be acceptable when its benefits and replacement costs are understood.
The decision
A practical test of independence is whether an alternative works with what you can take with you.
Sources and data
The charts use a fixed snapshot so this briefing remains reproducible. Sources may continue to change. The switching checklist is a proposed discussion framework, not a result validated by these studies.
- 01Model Context Protocol architecture ↗
MCP documentation
Hosts and clients connect to servers that expose tools, resources and prompts. MCP itself does not establish company ownership or portability.
- 02The plunging price of thought ↗
Epoch AI · Luke Emberson & David Roodman · 22 September 2026
Cost at a fixed level of benchmark performance; estimates and limitations.
- 03Inference-cost: code and data ↗
David Roodman · repository snapshot
Figure 1.csv supplies the historical quarterly decline rates used in our comparison.
- 04Epoch Capabilities Index ↗
Epoch AI · data snapshot 28 September 2026
ECI point estimates by release date. The displayed frontier is a running maximum, not a forecast.
- 05WikiSkill ↗
Tang et al. · August 2026 preprint · Table 1
Compiling Agent Experience into Persistent Knowledge for Skill Evolution.
- 06AI Coding Dictionary ↗
Matt Pocock
Model, harness, agent and environment. Definitions adapted for an executive audience.