Skip to content
FrankX.AI
AI ArchitectureAug 21, 202612 min read2,398 words

AI Architecture 2026: Four Decisions Hard to Reverse

TL;DR

Most AI system choices are reversible in an afternoon. Four are not. Where the model call goes sets your swap cost. The orchestration shape sets your failure mode. Where the trust boundary sits decides whether a retrieved document can act. Where a long run lives decides which platforms stay open to you. Everything else is a refactor.

Frank Riemer
FrankX
AI Architect & Independent Creator
Ex-Oracle AI Architect · Starlight & ACOS Systems
Most AI system decisions are cheap to change. Four are not: the vendor boundary, the orchestration shape, the trust boundary, and where a long run lives.
Reading Goal

Identify the four irreversible AI architecture decisions—vendor boundary, orchestration topology, trust boundaries, and execution state.

TL;DR: Prompt wording, chunk size, model choice, and framework are usually cheap to change. Four choices compound. Where the model call goes sets what a vendor swap costs. What shape the loop is sets how the system fails. Where the trust boundary sits decides whether retrieved content can trigger a side effect. Where a long run lives quietly eliminates platforms before you evaluate them. On 28 July 2026, the Model Context Protocol removed protocol-level sessions and the initialize handshake, moving the fourth decision for teams using that revision. This is the argument version. The full seven-plane reference lives here.

In architecture reviews, people inspect the boxes. Incidents arrive through the seams.

You may recognize the meeting. Someone draws model, retrieval, tools, and agent; connects them with arrows; and the room relaxes because the system finally looks understandable. Weeks later, the failure comes from a boundary the diagram treated as a line: provider-specific code scattered through the product, untrusted text reaching a tool, a loop with no durable state, or an action no one can reconstruct.

The seam changes the character of a request. That is where review time earns its keep.

A tactile seven-layer AI system instrument with product, orchestration, context, model and action flows wrapped by policy and evidence rails. An original architecture navigator inspects the stack.
Intelligence sits inside a governed system. Memory feeds context; evidence closes the loop.

Of everything in that stack, four choices are genuinely expensive to walk back. The rest is a refactor you can do on a quiet Thursday.

Which AI architecture decisions are actually hard to reverse?

DecisionWhat it really setsCost to change later
Where the model call goesYour swap cost when a provider raises prices, deprecates a model, or has a bad weekEvery call site, if you skipped the indirection
What shape the loop isYour failure mode: silent drift, context rot, or contradictory mergesThe control flow, plus every eval written against it
Where the trust boundary sitsWhether text a stranger wrote can cause a side effectA security review of every tool you already shipped
Where a long run livesWhich deployment platforms remain open to youA migration, because it is a runtime property, not a config flag

Prompt wording, chunk size, rerankers, and frameworks can often move on a quiet Thursday. Spend the serious review time on the four.

An AI architecture team routes work across several model and tool surfaces while specialist droids inspect provider boundaries, permissions, long-running state, and evidence.
The routing decision includes the provider boundary, the work shape, authority, durable state, and the evidence left behind.

Decision 1: Where does the model call go?

The cheapest version is openai.chat.completions.create() scattered through the codebase. It works immediately and it is the single most common thing I have to help teams undo.

Models change monthly. The durable question is whether exactly one place in your system knows a model provider's name.

If there is one place, you get four things without further work: routing by task, fallback when a provider degrades, a cache you can actually reason about, and a single line item for spend. If there are ninety places, you get none of them, and the day you need them is the day you cannot have them.

A governed model gateway takes one request through hard eligibility checks, selects from several provider engines using quality, latency, cost and availability evidence, and exposes an explicit fallback lane.
The gateway owns eligibility, routing, typed output validation, failover and degraded-mode disclosure.

This is the base plane for a reason. Everything above it inherits its properties. A system with no fallback path has no fallback path at every layer, no matter what the higher layers do. If you want the concrete version of this in a Next.js codebase, the Vercel AI SDK write-up covers the indirection without the abstraction tax.

Decision 2: What shape is the loop?

This is the decision most often made by accident. Someone builds a workflow, calls it an agent, and inherits neither the guarantees of a workflow nor the adaptability of an agent.

There are four common shapes, but they live on two independent axes: who chooses the next step, and how many executors coordinate. Picking one is three questions.

A two-axis architecture workbench separates code-owned from model-directed control and single executors from coordinated teams, with retrieval, tool and memory reach shown as independent multipliers.
Agency and coordination are separate decisions. Increase either only when measured uplift exceeds the additional control cost.

The bias should be toward Q1. A fixed workflow that you wrote down is easier to debug, evaluate, price, and explain to the person who owns the incident. Reach for a loop when the steps genuinely cannot be enumerated.

The expensive asset is the evaluation suite. A workflow is graded on whether each step produced the right output. A loop has to be graded on its trajectory: did it take a sensible path, or stumble into the right answer? Those require different harnesses. Switch shapes after you have a hundred graded examples and you may discard the harness with the control flow. Modern agentic systems architecture goes deeper on trajectory evaluation; the seven pillars piece covers what a production loop needs around it.

Decision 3: Where does the trust boundary sit?

Here is the sentence that reframes this for most people: text that came back from a tool call is not your text.

A retrieved document, an API response, a fetched web page, or an uploaded PDF was authored outside your trust boundary. If that text reaches the position in your context window where instructions live, it is an instruction. The model has no reliable way to tell the difference, and “ignore any instructions in the following document” is also just text. OWASP's prompt-injection guidance treats this as an application-level risk that retrieval and fine-tuning do not fully remove.

An untrusted model proposal crosses deterministic contract, identity, authorization, risk, approval, scoped credential, execution, verification and evidence gates before becoming a committed effect.
A trusted action corridor turns an untrusted proposal into a bounded, verified effect. Credentials never enter model context.

The fix is structural:

  1. Untrusted content never enters the instruction position. Fence it, label it, keep it in the data position.
  2. Side effects sit behind a gate. A write, message, payment, or merge needs a step that a document cannot perform on its own behalf.
  3. Tool scopes are narrow by default. A tool that can read a calendar and a tool that can read a calendar and send email are different risk objects, even though they are one line apart in your config.

The repair may be small. The audit grows with every tool already shipped against the old boundary. Draw that boundary before the tool surface expands, and it stays cheap. Zero-trust tool meshes covers the production shape of this.

Decision 4: Where does a long run live?

An agent loop that runs for eleven minutes creates a different deployment problem from a chat completion that returns in two seconds.

This is where teams discover their platform choice was load-bearing. A serverless function has a duration ceiling. A loop that outlives a request needs somewhere else to be: a queue and a worker, a durable execution primitive, a container that stays warm, or a machine you start and stop on purpose.

Nobody sits down and decides this. It gets decided by whatever the marketing site was optimized for, and it is discovered in week six.

A durable state-machine railway records intent, executes with an idempotency key, verifies and commits results, and branches into retry, reconciliation, compensation, waiting, cancellation or human repair.
A long run needs persistent state and explicit recovery paths. The workflow engine remembers; the model does not.

There is no universal platform winner. Managed platforms give you a primitive and absorb part of the operations; self-assembled clouds give you every primitive and take your Tuesdays. Both can be correct, depending on what the team wants to own. Vercel Workflows, for example, provides managed durable execution for long-running applications and agents. Queues, workers, containers, and other workflow engines create different ownership contracts.

Long-running work is a runtime property. A configuration flag cannot move it safely. Migration means moving the execution model and the state implied by that model.

What changed in mid-2026 that moves these decisions?

One thing, concretely, and it lands on decisions 3 and 4.

On 28 July 2026, the Model Context Protocol shipped a revision that made the protocol stateless. Reading the official changelog directly, the load-bearing removals are:

  • Protocol-level sessions are gone, along with the Mcp-Session-Id header. Servers that need state across calls now mint explicit handles and pass them as ordinary tool arguments.
  • The initialize handshake is gone. Every request carries its own protocol version and client capabilities in _meta. A new server/discover RPC advertises what a server supports, if you want to ask up front.
  • Stream resumability is gone. No Last-Event-ID, no message redelivery. A broken response stream loses the in-flight request, and the client has to re-issue it as a new request with a new ID.

Read those three together and the architectural consequence is clear: state that used to be implicit in a connection is now something you have to name and own. This adds work at the seam while removing connection state that made horizontal scaling and failure analysis harder.

The third bullet is the one that changes plans. If your design assumed a long-lived stream would survive a blip, it no longer does. Retries become your problem, which means idempotency becomes your problem, which means decision 4 just got more expensive to get wrong. The 2026 agentic hierarchy covers where MCP sits relative to skills and agents if that layering is still fuzzy.

How do you tell which decisions you already made by accident?

Four greps and one question. This takes about twenty minutes on a codebase you know.

  1. Grep for your model provider's SDK import. More than one module means decision 1 is deferred, not made.
  2. Find the loop's exit condition. If it lives in a prompt rather than in code, you have an unbounded loop with a polite request attached.
  3. Trace one retrieved document from the retriever to the context window. If you cannot point at the line where it becomes labelled data rather than plain text, decision 3 has been made for you.
  4. Find your longest production run. Use the longest run rather than the average. Then check the platform ceiling. If you do not know either number, decision 4 is unmade.

The question: what breaks first if your primary model provider is unavailable for four hours? If the answer is "everything," the four decisions are all still open, and they are open in the expensive direction.

FAQ

Q: Is a fixed workflow really better than an agent? Better at different things. A workflow is cheaper, more debuggable, and easier to evaluate. An agent adapts to inputs you did not anticipate. The failure is picking an agent for work you could have enumerated, because you then pay the agent's cost and get the workflow's ceiling. Start with the workflow and let a real failure push you to the loop.

Q: What is the trust boundary in an AI system? The line between text you authored and text that arrived from somewhere else. System prompts and operator instructions are trusted. Retrieved documents, tool results, fetched pages and uploaded files are not, regardless of how reputable the source looks. The boundary is architectural: untrusted content stays in the data position of the context window and never reaches the instruction position, and side effects sit behind an approval gate.

Q: Did MCP really remove sessions? Yes. The 28 July 2026 revision removed protocol-level sessions and the Mcp-Session-Id header, removed the initialize handshake, and removed SSE stream resumability. Servers needing cross-call state now mint explicit handles passed as ordinary tool arguments. The changelog in the specification repository is the authority.

Q: Which platform should I deploy an AI agent on? The one whose primitive matches how long your runs actually are. Short request-scoped work fits serverless functions and edge isolates. Work that outlives a request needs a durable execution primitive, a queue and worker, or a container that stays warm. Measure your longest real run before you choose, not your median.

Q: How many of these decisions can I defer? All four, briefly. Deferring decision 1 is cheapest and stays cheap longest. Deferring decision 3 gets expensive fastest, because the cost scales with the number of tools you have already shipped. If you defer one thing, do not let it be the trust boundary.

Where to go next

The four decisions are the argument. The reference is longer and lives on its own page: the seven planes of a production AI system, including what each plane owns, what the seam between two planes actually changes, and the failure modes that show up in the wrong plane from where they were caused.

For the operational layer, continue with the observability stack for multi-agent systems. Three of the four decisions above are only checkable when you can see what a run actually did.

Sources checked against primary documentation on 30 August 2026. Where a claim depends on a vendor page or a protocol revision, the linked source is the authority.

Method note: source review and editorial architecture were developed in ChatGPT Work Mode; repository implementation and verification used Codex; original scene layers were created with native OpenAI image generation. Exact architecture diagrams and claims remain deterministic and source-checkable.

Axi

Read on FrankX.AI — AI Architecture, Music & Creator Intelligence

Stay in the intelligence loop

Weekly field notes on AI systems, production patterns, and builder strategy.

Occasional FrankX field notes. Unsubscribe anytime. Privacy details.