AI Vendor Lock-In Will Look Different Than SaaS Lock-In

Why the next switching cost will live not only in data and applications, but in behavior, orchestration, evaluation, and operational learning

We already know what software lock-in looks like.

A company introduces a CRM, ERP platform, service-management system, HR suite, billing platform, or another major SaaS application. Over time, data accumulates. Processes adapt. Interfaces multiply. Custom configurations grow. Employees learn the platform. Partners specialize in it. Commercial commitments become larger.

Years later, replacing the system is technically possible, but economically painful.

None of this is new.

And artificial intelligence does not make these familiar forms of dependency disappear. AI platforms can still create proprietary APIs, contractual commitments, integration costs, data gravity, configuration complexity, and ecosystem dependency.

But AI introduces another layer that traditional software rarely had in the same form.

An enterprise can become dependent not only on where its data and processes live, but on how a particular combination of models, prompts, retrieval systems, agents, tools, policies, and runtime services behaves.

That distinction matters.

With traditional SaaS, the migration problem is often approximately:

Export the data. Reproduce the configuration. Replace the integrations. Retrain the users.


With a mature AI capability, the problem may increasingly become:

Reconstruct enough of the organization’s machine-readable operating logic to reproduce an acceptable business outcome on a different stack.


That is a different kind of switching cost.

In previous essays, I argued that usable enterprise data, architectural discipline, risk-based governance, AI readiness, and a serious understanding of agent-enabled work are prerequisites for scaling AI.
The next question is what happens after companies have built those capabilities.

What exactly are they becoming dependent on?

And, more importantly, which of those dependencies are they consciously willing to accept?

We Already Know What SaaS Lock-In Looks Like

Traditional SaaS dependency is relatively easy to describe.

There is data gravity. Years of transactions, metadata, attachments, audit history, permissions, relationships, workflow state, and configuration accumulate inside the platform. An export function does not necessarily make all of that meaning portable.

There is process dependency. Organizations gradually organize work around what the software can do. A process that once preceded the application becomes shaped by it.

There is configuration dependency. Forms, workflows, permissions, scripts, reports, business rules, extensions, and custom objects become part of the operating environment.

There is integration dependency. Other systems learn the vendor’s API semantics and data structures. What began as one application becomes a node with dozens or hundreds of connections.

And there are commercial and human dependencies. Enterprise agreements, minimum commitments, implementation partners, skills, certifications, habits, and internal expertise all create friction around change.

I discussed some of this in The Return of Architecture. The important point here is not to repeat the SaaS argument. It is to establish the baseline.

Traditional lock-in is frequently concentrated around data, configuration, process, integration, contracts, and skills.

AI retains all six. Then it adds a qualitatively different form of behavioral dependency: probabilistic behavior that is only partially specified and must be empirically evaluated.

AI Systems Contain Behavior, Not Just Functionality

Traditional enterprise software can be enormously complex, but most of its important behavior is deliberately specified.

If a billing rule says that a particular tariff applies under particular conditions, that rule can usually be located somewhere. It may be ugly. It may be buried in decades of custom code. But deterministic business logic has an important property: given the same state and the same inputs, we generally expect the same defined behavior.

Generative AI changes this relationship.

Large language model applications are probabilistic systems. OpenAI’s current model-optimization guidance explicitly states that LLM outputs are non-deterministic and that behavior changes across model snapshots and model families. It therefore recommends an iterative process of evaluation, prompting, testing, and tuning.

This means a company operating AI in production gradually learns things that do not exist neatly inside a classical requirements document.

Which model performs reliably for this particular task?

How should context be structured?

Which system instructions improve consistency?

When should retrieval run?

How many documents should be retrieved?

Which cases require reranking?

When does the model call the wrong tool?

Which instructions reduce that failure?

When must a human intervene?

Which failure rates are acceptable?

Which languages behave differently?

What happens when a conversation becomes long?

Where does the model become overly cautious?

Where does it become insufficiently cautious?

How should an agent recover after a failed tool call?

Over time, these answers become an operational asset.

And some of that asset can become coupled to the underlying model.

Current provider documentation gives unusually clear evidence of this. Anthropic’s migration documentation for newer Claude models does not merely describe renamed endpoints. It documents changes in reasoning behavior, safety handling, tokenization, prompting recommendations, tool behavior, and other characteristics that developers should re-baseline when migrating. Its documentation for Claude Opus 5 explicitly notes that the model can behave differently from its predecessor even when application code is unchanged.

That does not mean every model upgrade requires a rewrite.

It means something more subtle.

Behavioral compatibility becomes an additional dimension of portability.

From Data Portability to Behavioral Portability

“Behavioral portability” is not a formal industry standard I am invoking here. I use it as a management shorthand for a very practical question:

Can the business behavior we depend on be reproduced on another model or platform at an acceptable level of quality, risk, cost, and performance?

This is different from asking whether the assets can technically be copied.

A system prompt can be copied.

A JSON schema can be copied.

A tool definition can be copied.

A set of retrieval documents can be copied.

A workflow definition may be exportable.

But those facts alone do not tell you whether another model will select the same tools, interpret ambiguous instructions similarly, produce equally usable structured output, follow the same escalation rules, exhibit the same safety behavior, or perform equally well on the particular edge cases your organization has spent two years discovering.

Research on prompt sensitivity is nuanced. Some studies have found significant changes in model performance under relatively small variations in instructions, while more recent work has shown that parts of the apparent sensitivity can be exaggerated by underspecified prompts or overly rigid evaluation methods. That is exactly why the correct conclusion is not “prompts are non-portable.” It is that operational equivalence has to be tested rather than assumed.

A prompt library can therefore be a valuable enterprise asset.

But its value does not reside only in the text.

It resides in the tested relationship between that text, the surrounding context, the model, the tools, and the acceptance criteria.

A prompt can be copied in seconds. Reproducing the behavior it helped create may take considerably longer.

The Model May Not Be the Real Lock-In

Much of today’s portability discussion focuses on foundation models.

Can we replace Model A with Model B?

That is a useful question. But for many enterprise systems it will eventually become the smaller question.

Consider what surrounds the model in a production agent.

There may be workflow state, memory, tool definitions, delegated credentials, authorization policies, retrieval, business rules, approval steps, error handling, retries, human escalation, scheduling, logging, evaluation, cost controls, tracing, security controls, and integrations with enterprise applications.

Modern managed agent platforms are increasingly integrating precisely these capabilities. Amazon’s current AgentCore documentation, for example, describes runtime, memory, identity, gateway, observability, policy, tools, and evaluations as parts of its managed agent environment. Google’s current Gemini Enterprise Agent Platform similarly combines runtime, sessions, memory, identity, security, observability, evaluation, gateways, retrieval, and other agent services.

That integration can be extremely valuable.

A company does not have to build credential management, memory services, distributed tracing, policy enforcement, deployment infrastructure, and agent lifecycle management from scratch. Managed integration can dramatically reduce implementation effort and operational risk.

But integration also changes the object of dependency.

I find it useful to distinguish four layers.

Model lock-in is dependency on a particular foundation model and its characteristics.

Platform lock-in is dependency on the surrounding APIs, runtime, security, storage, management, and development services.

Agent lock-in is dependency on accumulated workflows, state, memory, tools, permissions, delegation rules, evaluations, and operating logic.

Ecosystem lock-in appears when AI becomes intertwined with the broader cloud, identity, productivity, data, observability, security, and development environment of one supplier.

This is not meant as a formal taxonomy. Its purpose is to expose an architectural mistake.

A company can have excellent model portability and still have terrible capability portability.

It may take one line of configuration to change the foundation model while requiring a major transformation program to leave the agent platform around it.

The model may be replaceable while the capability built around it is not.

RAG and Embeddings Create Derived Dependencies

Data ownership remains fundamental.

But AI adds an important distinction between source assets and derived AI assets.

Suppose an enterprise owns every document used by its retrieval system. It has retained the originals. The files are stored in open formats. Nothing is trapped inside the model provider.

That is good.

It is not the same as saying that the knowledge architecture built around those documents is portable.

A production RAG system may contain parsing rules, normalized documents, chunking strategies, chunk boundaries, metadata enrichment, access-control metadata, embedding models, vector indexes, hybrid search logic, rerankers, retrieval thresholds, query transformations, context assembly, prompt templates, and evaluation datasets.

These choices materially influence retrieval. Current Microsoft architecture guidance for RAG, for example, treats chunk cleaning, metadata enrichment, vectorization, and indexing as explicit design stages. Amazon Bedrock similarly documents configurable fixed, hierarchical, semantic, and multimodal chunking before content is embedded into a vector index.

This creates a useful management distinction:

Owning the source documents is necessary. It does not mean you automatically own a portable implementation of the knowledge system built from them.

Embeddings are an especially clear example.

An embedding is not simply a compressed copy of the source text. It is a representation created by a particular model inside a learned vector space. Separately trained embedding models generally do not produce vectors that can simply be substituted for one another. Research into cross-model embedding alignment exists precisely because separately trained spaces are often not directly interchangeable.

In practical systems, changing the embedding model commonly means regenerating embeddings and rebuilding the relevant index. Weaviate’s guidance explicitly warns that changing embedding models affects downstream results and can require re-embedding and re-indexing the data.

The management lesson is not that embeddings are a dangerous technology.

It is much more useful:

Derived AI artifacts can create migration work even when the enterprise retains complete ownership of the original data.

And migration is not finished when the new vectors have been generated. Retrieval quality must be reevaluated because a new representation can change which information is considered semantically close.

RAG therefore improves one important form of control by keeping proprietary knowledge outside the foundation model’s weights and maintaining it as an independently managed enterprise asset.

Customization Can Create Valuable Capability and Deeper Coupling

The same principle applies to model customization.

Fine-tuning, parameter-efficient techniques such as LoRA, proprietary customization APIs, distillation, or other model adaptation methods can create substantial business value. A customized model may perform a narrow domain task more consistently, more cheaply, or with less context.

But customization asks a second question alongside performance:

Where does the resulting capability actually live?

With LoRA, for example, adapter parameters are applied to a particular base model structure. Hugging Face’s PEFT documentation explicitly treats the base model as part of the adapter configuration and execution model.

In a hosted proprietary customization service, the dependency can be different again.

A particularly current illustration is OpenAI’s own fine-tuning lifecycle. As of August 2026, OpenAI’s documentation says it is winding down its current fine-tuning platform and states that existing fine-tuned models will remain available only until their underlying base models are deprecated.

That is not an argument against customization or against OpenAI. Product evolution is normal in a market moving this quickly.

It is an argument for asking the right questions before customization becomes strategically important.

Do we retain the training dataset?

Do we retain the evaluations?

Can the customization artifact itself be exported?

Is it meaningful without the original base model?

Could the capability be reconstructed elsewhere from the data we own?

What happens when the model family is retired?

The answers will differ by technology and provider.

That is precisely why “we own the data” is too incomplete as an exit strategy.

The Most Expensive Dependency May Be What You Learned

There is another form of lock-in that will rarely appear on a vendor invoice.

Operational knowledge.

Imagine an enterprise AI capability that has been in production for three years.

During that time, teams have discovered unusual failure modes. They have written exception handling. They have adjusted prompts. They have learned which model works best in German customer correspondence and which works better for code. They have built escalation rules. They know which retrieval threshold creates too many false positives. They have calibrated human review. They know which tool calls need extra validation. They have built incident procedures. They understand acceptable latency and cost. They have learned which outputs users distrust and which ones require extra explanation.

None of that knowledge was present on day one.

It was earned through operations.

OpenAI’s current model-optimization guidance effectively describes this type of continuous learning loop: establish evaluations, run representative inputs, inspect results, improve prompts or training data, and repeat as model behavior and requirements evolve.

This is why I suspect that one of the most consequential forms of AI lock-in will be organizational rather than contractual.

The most expensive AI dependency may not be the API. It may be everything the organization has learned about making that API reliable.

Some of that knowledge transfers beautifully.

Good evaluation methodology transfers. Domain understanding transfers. Well-designed business rules transfer. Strong datasets transfer. Mature incident management transfers.

Other knowledge may turn out to be highly model-specific.

That distinction is strategically important because it tells you what should remain an enterprise asset.

Evaluation Is Portability Infrastructure

This leads to one of the strongest control points available to an enterprise: evaluation.

A mature AI organization should have its own definition of “good.”

That definition might include golden test cases, domain-specific benchmarks, factuality requirements, expected classifications, accepted output structures, tool-selection tests, retrieval quality, safety checks, adversarial scenarios, human-review protocols, latency targets, cost thresholds, and business outcome measures.

The exact evaluation method will depend on the capability.

But the strategic principle is stable.

If only the vendor can tell you whether the AI system is good enough, you do not fully control the capability.

Evaluation assets provide a bridge between vendors.

They allow the company to ask not merely whether another model can accept the same API call, but whether it can meet the actual business requirement.

This is especially important because provider documentation itself recommends evals when models change. OpenAI says model behavior can change across snapshots and families and recommends systematic evaluation. Anthropic’s migration guidance similarly tells customers to revisit behavior, prompts, token budgets, and workload performance when moving between models.

There is an almost perfect illustration of the principle in the current tooling landscape: OpenAI continues to emphasize evaluations as an essential discipline while simultaneously deprecating a particular generation of its hosted Evals platform.

The methodology survives the tool.

That is exactly how enterprises should think.

Portability is not merely the ability to move technology. It is the ability to independently prove that the replacement still fulfills the business requirement.

Abstraction Helps. But Dependency Often Moves Upward.

The obvious architectural response is abstraction.

Create an AI gateway.

Standardize model APIs.

Put a common interface around tool calling.

Route requests across multiple providers.

Separate the application from the model.

All of these can be sensible.

For some workloads, model substitution really can become relatively easy.

But abstraction is not free.

Every abstraction layer has to decide which differences to expose and which to hide. If it hides too little, every application still understands every vendor. If it hides too much, the organization may lose access to precisely the differentiated capabilities it was paying for.

A lowest-common-denominator interface can make portability excellent while making the product mediocre.

An abstraction platform can also become a new control plane containing routing logic, credentials, rate limits, prompt management, evaluation, logging, policy enforcement, fallback rules, and cost management.

At that point the question has changed from:

Are we locked into the model?


to:

How replaceable is the layer we built to avoid being locked into the model?


Multi-model architecture has the same limitation.

Using three models may reduce dependency on any one foundation model. But if all three run through one proprietary orchestration platform, one vector database, one identity environment, one agent runtime, or one evaluation service, the dependency has not disappeared.

It has moved.

The abstraction layer you build to avoid one vendor can become your next platform dependency.

That does not make abstraction a mistake.

It means abstraction should be applied selectively, where the expected option value exceeds its architectural and operational cost.

Open Standards Reduce Switching Friction, Not Behavioral Differences

There is genuine progress on interoperability.

OpenAPI remains a mature, vendor-neutral way to describe HTTP APIs and has broad tooling support.

For agent-to-tool integration, the Model Context Protocol has developed quickly. The current MCP specification, released on July 28, 2026, provides standardized mechanisms around tools, resources, capability discovery, and HTTP authorization. Its current authorization design builds on established web standards including OAuth-related mechanisms.

For agent-to-agent communication, A2A provides an open specification for agent discovery, communication, and task delegation across heterogeneous agent implementations.

And OpenTelemetry is extending vendor-neutral telemetry into generative AI, including model operations, token usage, retrieval, tool activity, and other agent traces. OpenTelemetry itself is now a graduated CNCF project, although its GenAI-specific conventions are still evolving rapidly.

This is significant progress.

But standards need to be understood at the layer they actually standardize.

MCP can make a tool interface more reusable.

A2A can make agent communication less proprietary.

OpenTelemetry can reduce observability coupling.

OpenAPI can describe an HTTP interface.

None of them guarantee that Model B will select the same tool as Model A, that two agents will plan identically, that a replacement retrieval stack will rank the same passages, or that a safety model will make the same decisions.

There is another useful warning here. Open standards themselves evolve. The July 2026 MCP release included substantial protocol changes and a formal migration cycle for implementers.

That is normal.

It simply illustrates the broader point:

Interface standardization can reduce integration lock-in. It cannot eliminate semantic or behavioral dependency.

API compatibility is not behavioral compatibility.

Identity and Observability Can Become Control Points

Agentic systems add another enterprise dependency that simple chatbot deployments did not expose as clearly.

Authority.

An agent is increasingly expected not only to generate text but to act.

It may read a customer record, update a ticket, issue a refund recommendation, create a purchase request, send a message, retrieve confidential data, run code, or invoke another agent.

That requires identity, authorization, delegated authority, credentials, policies, approval, and auditability.

Modern agent platforms are already building dedicated agent-identity mechanisms. AWS describes AgentCore Identity as an authentication, authorization, and credential-management service for agents acting across AWS and third-party systems. Google’s Agent Identity similarly provides per-agent permissions and auditability when agents act on users’ behalf.

This is extremely useful functionality.

But if an enterprise embeds its entire delegated-authority model inside one proprietary agent platform, a future migration becomes more than an AI migration.

It becomes an identity and security migration.

The same applies to observability.

Production AI systems generate traces that can include prompts, responses, retrieved documents, tool calls, token consumption, latency, policy decisions, errors, evaluations, user feedback, and agent execution paths. That operational history can become part of the organization’s knowledge about how its AI capability actually behaves.

A sensible control question is therefore not simply:

Can we export the logs?


It is:

Which operational evidence must remain available to us independently of the runtime producing it?


This is one reason standards such as OpenTelemetry matter strategically, even though they do not create complete portability.

The Hyperscaler Question Is About Coupling, Not Avoidance

It would be easy to turn this argument into an anti-hyperscaler position.

That would be a mistake.

The integrated stack is the point.

Enterprises buy cloud ecosystems precisely because they can combine infrastructure, security, identity, data, networking, AI models, managed runtimes, observability, policy, vector search, integration, and support.

That can create huge operational advantages.

A heavily integrated managed agent platform may be faster to implement, easier to secure, more observable, and more resilient than an internally assembled collection of twenty loosely coupled components.

The strategic question is therefore not:

Should we avoid hyperscalers?


It is:

Which layers of our AI operating model are we comfortable coupling to one ecosystem?


A low-risk employee assistant may justify very deep coupling because portability would have little economic value.

A revenue-critical customer capability, regulated decision-support process, essential knowledge service, or strategically differentiating product may deserve a very different answer.

The right level of portability is contextual.

Open Models Solve Some Lock-In Problems, Not All of Them

The same discipline is necessary when discussing “open source AI.”

Terminology is often used far too loosely.

The Open Source Initiative’s Open Source AI Definition distinguishes a fully open-source AI system from the narrower concept of open weights. Access to model weights alone does not necessarily provide all of the information and freedoms that OSI associates with Open Source AI.

That distinction is operationally useful even if one does not want to enter the broader debate around the definition.

Open or downloadable weights can reduce important dependencies.

They may allow an enterprise to select its own infrastructure, operate a model in different environments, retain more control over deployment, or continue using a model independently of a hosted inference endpoint, subject to the applicable license.

That can be strategically valuable.

But it does not remove infrastructure, accelerator, serving, security, data-pipeline, retrieval, evaluation, observability, skills, or operational dependencies.

In fact, some responsibility moves in the opposite direction.

When a managed service is removed, somebody has to operate what the managed service used to provide.

There is no contradiction here.

Open models can increase control over one layer while increasing enterprise responsibility for another.

The same is true of proprietary models in reverse: less direct control can be a rational price for superior capability, faster delivery, lower operational burden, or stronger managed services.

That is why the useful debate is not “open versus closed.”

It is control versus dependency, layer by layer.

Contracts, Lifecycles, and Economics Can Override Architecture

A technically portable architecture can still be commercially difficult to leave.

An enterprise AI contract may need to address matters such as customer data, retention, deletion, output rights, customization assets, telemetry, usage commitments, termination, transition periods, model changes, pricing changes, and post-termination access.

The precise legal position depends on the contract, jurisdiction, sector, service, and provider. This is not a legal argument.

It is an architectural one:

A contract can promise data portability while the architecture makes migration expensive. An architecture can be technically portable while the commercial arrangement makes exit unattractive.

Both disciplines have to align.

Regulated sectors already understand this logic. The EU’s Digital Operational Resilience Act, for example, explicitly addresses ICT third-party risk and requires financial entities to consider exit strategies and transition arrangements for relevant ICT dependencies.

AI also makes provider lifecycle risk more visible.

OpenAI currently publishes defined deprecation notices and migration paths for API models and services. Anthropic similarly maintains explicit model lifecycle and deprecation documentation.

Again, there is nothing unreasonable about this.

Foundation models are improving unusually quickly. Providers cannot maintain every historical model indefinitely.

But a company should know what happens when a production process depends on something that is retired.

Can the replacement pass the evaluation suite?

Do prompts need revision?

Do token economics change?

Does latency change?

Do safety characteristics change?

Do tools behave differently?

How long is the migration window?

This is where model lifecycle management becomes business continuity management.

Economics add another layer.

AI consumption is often more variable than classic seat-based SaaS. Current provider pricing models can distinguish input and output tokens, cached input, tools, modalities, storage, searches, or other runtime services.

Prices can decline. New models can be substantially more efficient. Architectures can become cheaper over time.

But the reverse strategic point still stands.

If the economics of a business process rely heavily on one particular model’s price/performance profile, a pricing or model change can alter the business case without any change in the process itself.

Commercial portability therefore belongs in the architecture conversation.

Zero Lock-In Is the Wrong Goal

At this point there is an obvious overreaction available:

Abstract everything.

Never use a proprietary feature.

Maintain multiple providers everywhere.

Self-host every important model.

Keep two complete agent runtimes operational.

Avoid vendor-specific services.

Build every control layer internally.

That would be just as simplistic as ignoring lock-in completely.

Portability has a cost.

It creates engineering work, testing work, operational complexity, additional interfaces, duplicate expertise, and sometimes inferior access to vendor innovation.

A company can spend so much protecting itself from a hypothetical future migration that it destroys the economic advantage of the technology it is buying today.

That is not strategic independence.

It is expensive indecision.

The objective should therefore not be zero lock-in.

It should be conscious dependency.

For every meaningful dependency, the organization should understand what it is giving up, what it receives in return, what the exit would cost, and whether that tradeoff is rational given the business criticality of the capability.

This is where I would use the idea of reversibility rather than perfect portability.

Portability asks whether something can move.

Reversibility asks whether the organization can credibly change course.

Those are not always the same thing.

Own the Strategic Control Points

An enterprise does not need to own every layer of the AI stack.

But it should deliberately identify the layers where loss of control would materially damage its ability to differentiate, negotiate, comply, operate, innovate, or switch.

For many organizations, the strongest candidates will be source data, business semantics, critical business rules, evaluation criteria, identity and authorization policy, important integration contracts, and domain-specific operational knowledge.

Orchestration may be strategically important for one company and entirely commoditized for another.

The foundation model may be a control point for one product and a replaceable utility for another.

The correct answer depends on the capability.

A simple control matrix can make that discussion explicit:

LayerBusiness ImportanceSwitching DifficultyWhat We ControlCredible Exit
Foundation modelLow / Med / HighLow / Med / Highprompts, evals, routing?tested replacement?
Knowledge / RAGLow / Med / HighLow / Med / Highsource data, metadata, chunking?rebuild index?
Agent / orchestrationLow / Med / HighLow / Med / Highworkflow logic, state, tools?export or reconstruct?
Identity / authorizationLow / Med / HighLow / Med / Highpolicies, roles, credentials?transferable trust model?
EvaluationLow / Med / HighLow / Med / Hightest data and criteria?runnable independently?
ObservabilityLow / Med / HighLow / Med / Hightraces and history?independent telemetry?
Commercial modelLow / Med / HighLow / Med / Highcontract and commitments?termination and transition?

The purpose is not to make every cell green.

A high switching difficulty may be entirely acceptable.

What matters is that it is known.

A dependency deliberately accepted in exchange for substantial value is a strategy. A dependency nobody understood until the vendor relationship changed is an accident.

Every Critical AI Capability Needs an Exit Test

For a strategically important AI capability, an architecture board or executive committee should be able to answer a relatively small number of questions:

  1. What exactly are we renting? Which model, platform, runtime, agent, data, security, observability, and managed capabilities come from the supplier?
  2. What do we retain independently? Source data, business semantics, prompts, rules, evaluation datasets, policies, telemetry, training data, and operational knowledge?
  3. Which assets are derived from the current AI stack? Embeddings, indexes, memory, generated metadata, customization artifacts, workflow state, or vendor-specific evaluations?
  4. Which parts can actually be exported and reused, rather than merely downloaded?
  5. Can another model or platform reproduce the business behavior we require?
  6. Do we possess an independent evaluation suite capable of proving that?
  7. Which integrations, tool interfaces, identities, permissions, and policies are vendor-specific?
  8. What would have to be regenerated, reconfigured, revalidated, or completely rebuilt?
  9. What do our contracts permit us to retain, transfer, delete, or continue using after termination?
  10. What would a migration realistically cost in money, time, operational risk, and temporary capability loss?
  11. How much disruption could the business tolerate given the criticality of the capability?
  12. What measurable advantage are we receiving in return for accepting the dependency?

An enterprise does not need a permanently running second provider for every system.

For most applications that would be wasteful.

But for strategically critical capabilities, management should at least know what departure would involve before departure becomes urgent.

A vendor exit strategy that exists only in a contract is not an exit strategy.

Vendor Lock-In Is a Management Decision

None of this belongs exclusively to enterprise architecture.

Business leadership determines how critical the capability is.

Architecture identifies dependencies and strategic control points.

AI and technology teams understand the actual technical coupling.

Data leadership protects source assets, semantics, and derived knowledge.

Information security determines how identity, authorization, credentials, and delegated authority are controlled.

Procurement understands commitments and negotiating leverage.

Legal understands rights, obligations, termination, IP, data, and transition provisions.

Finance understands the economics of duplicated capability and migration.

Risk and compliance determine what resilience is proportionate to the consequence of failure.

These are not sequential tasks.

They are one decision viewed from different angles.

Vendor lock-in is not a procurement problem that architecture discovers later. It is an architectural and strategic decision that procurement must help make explicit.

That is particularly important in AI because dependency may become harder to see.

The database still belongs to you.

The prompt is sitting in Git.

The API is standardized.

The model can theoretically be changed.

And yet the capability may be deeply dependent on a particular combination of behavior, retrieval, memory, tools, evaluation, authorization, runtime services, and operational learning.

On paper, everything looks portable.

Operationally, very little may be.

Conclusion: Choose Your Dependencies Before They Choose You

Every generation of enterprise technology has created dependency.

Mainframes did.

Databases did.

ERP systems did.

Cloud platforms did.

SaaS did.

AI will not be different in that respect.

What will be different is the shape of the dependency.

Traditional lock-in often accumulated around data, configuration, integrations, processes, contracts, and skills.

AI adds another set of assets: model behavior, prompt-model interactions, derived representations, retrieval architecture, agent logic, tool behavior, memory, evaluations, telemetry, permissions, and everything an organization learns while making probabilistic systems reliable enough for real business operations.

That is why owning your data, while essential, is no longer sufficient.

Using an open API is not sufficient.

Being able to switch model IDs is not sufficient.

Using multiple models is not sufficient.

Using open weights is not sufficient.

And abstracting every vendor-specific feature is neither sufficient nor necessarily desirable.

The correct management objective is not technological purity.

It is deliberate architectural control.

Know what you own.

Know what you rent.

Know what you cannot easily replace.

Know what you would have to relearn.

Know what advantage you receive for accepting that dependency.

And invest in reversibility in proportion to the consequences of being unable to change course.

Because the strategic question is not whether your company will have AI dependencies.

It will.

The strategic question is whether those dependencies were consciously chosen, economically justified, architecturally understood, and reversible where it matters.

← Previous essay

Next essay →

Join the discussion

Your email address will not be published. Required fields are marked *

Keep reading

Newsletter

The occasional dispatch.

One considered email when something is worth your attention — new essays, conversations, and the signals behind them.

Used only for the dispatch. No tracking. Unsubscribe anytime.