At the end of August, as we started the migration, I opened the Helm chart our MCP servers deploy with, written in the spring, and read this, two lines above the replica count:
# -- How many replicas of the (runner + edge) pod. Stateless → scale freely.
replicas: 2
Every server under that comment speaks MCP over Streamable HTTP with a session per client, and the Kubernetes Service in front of each one has no affinity. Two replicas would have split every session across two pods, and half of every client’s calls would have failed with Session not found. We got lucky. Every deployed copy overrides that value to one replica, so nothing ever broke. The comment was still wrong, and it was wrong in a way I didn’t have words for until the new spec landed.
On July 28 the Model Context Protocol shipped revision 2026-07-28. It deletes sessions from the protocol. Before migrating anything I audited every place NimbleBrain touches MCP. We’ve since moved every MCP server we run to the new revision, and this audit is what told us where to start. If you run MCP in production on either side of the wire, the same audit is waiting for you.
What MCP 2026-07-28 removes
I read the changelog with our stack open beside it. If you run MCP over HTTP, ten changes matter.
Sessions are gone. No Mcp-Session-Id header, no session to mint, no DELETE to end one. tools/list, resources/list, and prompts/list may not vary per connection anymore. If your server needs state across calls, it mints a handle and the client passes it back as an ordinary tool argument. The proposal that made this change (SEP-2567) calls these explicit state handles: create_counter() returns a counter_id, and increase_counter(counter_id) uses it.
The handshake is gone. No initialize, no notifications/initialized. Every request carries its own protocol version and client capabilities in _meta. Servers must implement a new server/discover method so clients can ask up front what they support. A version mismatch gets an UnsupportedProtocolVersionError listing the versions the server does speak, and the client retries.
Servers can no longer call clients. Elicitation, sampling, and roots used to be server-initiated JSON-RPC requests riding an SSE stream. Now they’re embedded in results. A tool that needs user input returns resultType: "input_required" with an inputRequests map and an opaque requestState blob. The client gathers the answers and retries the original call, under a new request id, with inputResponses and the exact requestState it received. The spec names this Multi Round-Trip Requests. The server holds nothing between the two calls; whatever it needs, it encodes into requestState, and it must integrity-protect that blob if it affects authorization or business logic.
The GET endpoint is gone. There is no standalone server-to-client stream. If a client wants tool-list changes or resource-update notifications, it opens a long-lived POST called subscriptions/listen and names the notification types it wants. Progress and log notifications for a specific request still flow on that request’s own response stream.
Resumability is gone. No Last-Event-ID, no event redelivery. A dropped response stream loses the request and the client re-issues it with a new id. Closing the stream is the cancellation signal.
Tasks left the core protocol. The 2025-11-25 tasks utility is now an extension, io.modelcontextprotocol/tasks, and it was redesigned on the way out: polling via tasks/get, client input via tasks/update, no tasks/result, no tasks/list.
Three RPCs were deleted. ping, logging/setLevel, and notifications/roots/list_changed. Log level is now set per request with io.modelcontextprotocol/logLevel in _meta, and a server must not send log notifications for a request that didn’t ask for them.
HTTP headers are mandatory. Every POST carries MCP-Protocol-Version, Mcp-Method, and, for tool calls, resource reads, and prompt gets, Mcp-Name. A tool can annotate a parameter with x-mcp-header and clients mirror the value into an Mcp-Param-* header. Servers must reject a request whose headers disagree with its body. The point is that a load balancer can route and rate-limit MCP traffic without parsing JSON-RPC.
List and read results carry cache hints. ttlMs and cacheScope are required on tools/list, resources/list, resources/templates/list, prompts/list, and resources/read. Combined with the rule that lists can’t vary per connection, a client can cache a server’s tool list at the granularity of (deployment, credentials) instead of (connection).
Roots, Sampling, and Logging are deprecated. They still work for at least twelve months, but new implementations shouldn’t adopt them. The suggested replacements are tool parameters, resource URIs, or server configuration for roots, direct provider APIs for sampling, and stderr or OpenTelemetry for logs.
A tool call, before and after. Same tool, same arguments.
POST /mcp
MCP-Protocol-Version: 2025-11-25
Mcp-Session-Id: 6f1c9e2a-… ← minted by the server at initialize
{ "jsonrpc": "2.0", "id": 7, "method": "tools/call",
"params": { "name": "get_contact", "arguments": { "contact_id": "c_42" } } }
POST /mcp
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: get_contact
{ "jsonrpc": "2.0", "id": 7, "method": "tools/call",
"params": {
"name": "get_contact",
"arguments": { "contact_id": "c_42" },
"_meta": {
"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": { "name": "nimblebrain", "version": "…" },
"io.modelcontextprotocol/clientCapabilities": {}
} } }
The second request is self-contained. A server that has never seen this client can answer it. A server that just restarted can answer it. The pod next door can answer it. Every other line in the changelog follows from that.
What does stateless mean for an MCP server?
That chart comment wasn’t a typo. I wrote it, and I meant “stateless” the way most of us use it: the pod holds no data. Everything durable lives in a data store, so you can kill any pod and lose nothing. By that meaning, every server under the comment was stateless, and the comment was true.
The transport underneath it was not. The Python SDK we run keeps a session object per Mcp-Session-Id in process memory. Route a call for that session to a different pod and the pod has never heard of it. So “stateless, scale freely” was true of the data layer and false of the transport, and nobody noticed because the two meanings had never needed to be separated. The old spec let a server be stateless in the sense that matters for correctness and stateful in the sense that matters for scaling, and it gave you no vocabulary for the difference.
2026-07-28 collapses the second meaning into the first. There is no transport state to be stateful about. If your pods hold no data, they can now scale the way the comment promised.
Where MCP sessions hide in a real stack
We sit on every side of the protocol at once, which made this a bigger audit than I expected. The runtime is an MCP client to every tool server a tenant installs. It is an MCP server to its own web client and to any external host that connects. The web client is an MCP Apps host that renders interactive UIs shipped by those servers. And we operate a fleet of MCP servers of our own behind an edge proxy: six on the day of the audit, eight by the time we finished. Four surfaces, one protocol, and I wanted to know where the session actually lived in each one.
Everything below is the state on September 1, before any migration work.
| Surface | What it runs | Where the session lives |
|---|---|---|
| Runtime as client | TypeScript SDK 1.29, Streamable HTTP | One connection per (workspace, connector), held for the life of the process. Session loss detected by matching the string Session not found in error messages. |
| Runtime as server | Same SDK, one endpoint at /mcp | An in-memory LRU of live transports keyed by Mcp-Session-Id, an identity-bound session registry with an 8-hour TTL, and a per-session in-memory task table. |
| Web client as MCP Apps host | Same SDK plus the MCP Apps SDK | One transport singleton per browser tab. A retry wrapper for session loss whose own comment admits the recovered session has an empty task table. |
| Fleet servers | Python, FastMCP 3.4 | The SDK’s default session-per-client transport. Identity arrives per request in headers from the edge. All data in Postgres. |
Some of what I found was reassuring.
Our fleet servers already met most of the new rules without trying. Identity comes from per-request headers set by the edge proxy after it verifies the tenant’s token, not from anything the session remembers. Data lives in Postgres behind row-level security keyed on that identity. The People server takes a contact_id on every call that touches a contact, and the Tasks server takes a task_id. That is the explicit-handle pattern. We never wrote it down as a pattern because it was just how you write a CRUD tool. The only per-session state in the entire fleet was the transport session itself.
The runtime’s /mcp server already returns 405 Method Not Allowed on GET. We never pushed anything down the server-to-client stream, and we said so in a comment above the handler. The new spec deletes the GET endpoint outright. That one was free.
Our edge proxy never parses JSON-RPC. It verifies a token, rewrites identity headers, strips anything a tenant shouldn’t be able to set, and streams the response back unbuffered. The new mandatory headers, Mcp-Method and Mcp-Name, are what a proxy like that wants and could never have before.
The rest of what I found was debt.
Why sessions caused our MCP debt
In June we wrote a design note about the runtime’s remote MCP connection lifecycle. The problem it documented: the runtime holds one open connection to every remote tool server for the life of the process, with no idle timeout and no eviction. Our deploy runbook covers what that costs. When a fleet server rolls a pod, every tenant runtime with that server installed keeps a session to a pod that no longer exists. Their next call fails. The interactive UI for that server fails to load. The runtime recovers on its own most of the time now, and the runbook’s last resort is still “bounce the runtime.”
The design note names the root cause in one sentence: the connection is transport, not capability. The runtime caches each server’s tool list on the live connection object, so it can’t close the connection without losing the tool list, so it can’t close the connection. The note makes decoupling tool metadata from the transport the next step, and defers idle eviction until that lands.
The tool list was tied to the connection because the old spec allowed a server to return a different tools/list per session. The SDK couldn’t know whether a given server did that, so the only safe cache key was the session, and the session lived on the connection. The proposal that removed sessions (SEP-2567) makes exactly this argument: the mere possibility of session-scoped lists forces every client to refetch per connection, even though almost no server varies them.
Under 2026-07-28 that possibility is gone by rule. A tool list is a property of the deployment and the credentials, it carries its own ttlMs, and any pod can serve it. The fix we planned in June is the protocol’s default in July. The hold-forever policy, the string-match on Session not found, the three-stage retry schedule, the pod-roll runbook, the empty task table after a browser reconnect: every one of these traces to Mcp-Session-Id, and every one of them stops being necessary when the header stops existing.
That is the thesis of this series: the migration deletes the reasons the hold-forever design existed. Swapping the headers is the smallest part of it.
MCP SDK support for 2026-07-28 (October 2026)
The spec was final on July 28. The tooling wasn’t, and the gaps decided the order we migrated in. Here’s where each one stood when we started, and where it stands today.
The TypeScript SDK split in half. v2 shipped the day before the spec as separate packages: @modelcontextprotocol/server, @modelcontextprotocol/client, @modelcontextprotocol/core, middleware for Node, Express, Fastify, and Hono, a codemod, and a server-legacy package. The monolithic @modelcontextprotocol/sdk that every existing project pins carries on as a 1.x maintenance line on the 2025 protocol. Even on v2, nothing speaks 2026-07-28 by default. A client opts in with versionNegotiation: { mode: 'auto' }, which probes with server/discover and falls back to initialize against an older server. A server opts in by hosting through createMcpHandler(factory), which serves 2026-07-28 and, by default, answers 2025-era requests statelessly from the same endpoint.
Tasks fell out of core, and the SDKs are still catching up. At launch neither official SDK served the new tasks extension. The TypeScript SDK’s v2.3.0 (October 2) stopped refusing hand-registered tasks/get and tasks/cancel handlers on 2026-07-28 connections, but full support for the extension is still an open issue. The Python SDK lists the tasks extension (SEP-2663) as a known gap in its 2.2.0 release notes, and its roadmap still defers it. FastMCP 4, released August 31 on top of the Python SDK v2, ships an implementation as a separate fastmcp-tasks package. Our long-running tools were built on the 2025-11-25 tasks utility, so we carry a small client of our own until the SDKs catch up.
The MCP Apps SDK lagged by six weeks. @modelcontextprotocol/ext-apps peer-depended on the old monolithic package until its v2 migration merged on September 8. A host that moved earlier carried both SDK lines in one bundle. The apps protocol itself rides postMessage between iframe and host, so session removal never touched it. The wart was packaging.
Extensions that need a server-to-client channel are dead on the new era. We have one. When a tool needs a file that lives on the host side, the tool server asks for it mid-call over the live connection. That request has no channel to ride on anymore. The spec’s guidance for the equivalent built-in feature, roots, is to pass the data as tool parameters. So on 2026-07-28 our tools take the content inline, and the extension stays on the 2025 era until the spec grows a hook for it (there’s an open proposal upstream).
Logs go quiet. On the new era a server must not emit log notifications for a request that didn’t include io.modelcontextprotocol/logLevel in _meta. The SDK client doesn’t attach one by default. Migrate a server and its handler logs vanish from the client until you ask for them. The first time it happens you’ll think something broke.
Auth never picks an era, in the TypeScript SDK at least. The spec says a 4xx without a modern error body means fall back to initialize. The TypeScript SDK’s auto mode treats a 401 or 403 on the probe as an auth error instead: the client completes auth and probes again, rather than deciding the server is legacy.
How to order an MCP migration
Servers go dual-era first, because a server that speaks both protocols works under clients that speak either. Clients go modern second, once every server they talk to can answer them. Legacy gets retired last, on your own schedule. The spec sets no removal date for the initialize-based versions (the twelve-month window covers Roots, Sampling, and Logging), and the clients you don’t control will be slow.
So the fleet went first, from FastMCP 3 to FastMCP 4, which negotiates per connection. The first server took a pin change and a new lockfile. All 148 of its tests passed, and the same process answered a 2026-07-28 server/discover with no session and a 2025-11-25 initialize that minted one. Its tool list came back with ttlMs: 0 and cacheScope: "private", the SDK’s conservative defaults, so set real cache hints when you get there.
By early October all of it was in production: every one of our servers spoke both eras, and the runtime had moved to the v2 packages with per-connection negotiation. Every server we run now negotiates 2026-07-28. Servers run by other companies will move on their own schedule, so the runtime speaks both eras and picks per server. The next post in this series is the move itself: what broke loudly, and the four breaks nothing reported.
Check your own deployment
If you run an MCP server behind a load balancer, grep your deployment config for the word “stateless” and ask which meaning the author had in mind. If you run an MCP client that holds connections open, find the comment that explains why, and check whether the reason was the session. I’d bet it was.
The protocol just removed the thing most of that code was working around. Now would be a good time to find out how much of it you can delete.
This series started as my contribution to the Agentic AI Foundation’s MCP 7-28 education campaign. I’m an AAIF ambassador, and I’ll be speaking about the MCP Apps host side of this migration at AGNTCon + MCPCon North America in San Jose on October 22.
Share this post