Agent Router v1.2.x
v1.2.0
✨ New Features
Agent Router
Envoy AI Gateway is now Agent Router, an Agentic AI Foundation project with the same code, maintainers, and Apache 2.0 license. The website moved to theagentrouter.ai and the repository to theagentrouter/agent-router; old links redirect. Your manifests, Helm values, and automation keep working unchanged.
MCP Gateway
prefixMode: NeverClients that hardcode tool names, such as interactive MCP Apps, no longer have to see the <backend>__ prefix. Set prefixMode: Never on the route or on individual backendRefs. A Never-mode backend must list its tools in toolSelector.include, and the controller marks the route NotAccepted if two Never-mode backends expose the same name. Prompts are exposed bare only when they are listed in the new promptSelector.include. Resources always keep the prefix. The default is still Always.
injectionPolicy: IfNotPresentLet users bring their own upstream token, for example a personal GitHub PAT forwarded with forwardHeaders, and fall back to a shared service key when they don't. With IfNotPresent, the configured API key is injected only when the target header is missing. This applies to header injection only.
mergeTypeThe SecurityPolicy and BackendTrafficPolicy that Agent Router generates for an MCPRoute used to replace any Gateway- or listener-level policy. Set securityPolicy.mergeType or backendTrafficPolicy.mergeType to StrategicMerge or JSONMerge to keep a Gateway-wide rate limit or ext-auth policy in force alongside MCP OAuth.
oauth.authorizationServerMetadataUrl points the controller at the exact RFC 8414 metadata document when the issuer URL doesn't lead to it. The controller fetches this document while reconciling, so it needs network access to the identity provider. If the fetch fails, reconciliation fails.
The initialize response now negotiates the protocol version with the client and backends instead of always answering 2025-06-18. Merged list responses also carry ttlMs and cacheScope caching hints. Support for the stateless 2026-07-28 MCP specification is in progress and not yet enabled.
Providers & API Compatibility
AWSOpenAI)Send OpenAI-format /v1/chat/completions and /v1/responses requests to Bedrock Runtime's OpenAI-compatible /openai/v1 endpoint without a translation hop. Requests are signed with SigV4 through an AWS BackendSecurityPolicy. The prefix defaults to openai/v1. Other endpoints return 422.
TypeSafe)Route TypeSafe's Jev decision model through the gateway at /typesafe/v1/systemone with API-key auth. Request bodies pass through unchanged apart from a model name override, and token usage feeds the standard metrics and llmRequestCosts. Streaming is not supported.
anthropic-beta filteringOne anthropic-beta value that a provider rejects no longer fails the whole request. AIServiceBackend.spec.headerValueFilters drops listed values (Denylist, the default) or keeps only listed values (Allowlist). It applies to /v1/messages traffic to GCPAnthropic and AWSAnthropic backends. Bedrock also accepts four more beta flags, including thinking-token-count-2026-05-13.
Claude Opus 5 and 5.5 accept reasoning_effort on Vertex AI and Bedrock, and structured output (response_format) on Vertex AI. Chat-completion tools forward eager_input_streaming to Claude. Gemini on Vertex AI accepts vLLM-style structured_outputs (json, regex, choice).
cache_write_tokensPrompt-cache writes reported as cache_write_tokens by OpenAI Chat Completions and the Responses API now count toward cache-creation token metrics and costs. Translated Claude and Bedrock responses report the same value.
Model Catalog & Observability
/v1/modelsSet excludeFromModelsEndpoint: true on an AIGatewayRoute rule to keep internal or backwards-compatible model aliases out of /v1/models and /anthropic/v1/models. Requests to those models still route normally.
All four gen_ai.* metrics now carry a gen_ai.backend attribute (gen_ai_backend in Prometheus) set to the backend's namespace/name. This lets you compare latency, time to first token, and token usage across providers serving the same model.
Inference & Operations
An InferencePool without endpointPickerRef is now reported as Accepted=False (EndpointPickerRefMissing) instead of breaking routing. appProtocol: kubernetes.io/h2c is honored for cleartext HTTP/2 model servers.
The controller reports Ready only after its caches sync, so Envoy Gateway no longer calls the extension server with an empty view of your resources during a rollout. The chart adds a startup probe and a controller.priorityClassName value.
🔗 API Updates
AIGatewayRouteRule.excludeFromModelsEndpoint: Optional boolean (defaultfalse) in v1alpha1 and v1beta1.AIServiceBackend.spec.headerValueFilters: Optional list (max 16, one per header) of{name, mode, values}.modeisDenylist(default) orAllowlist;valuesholds up to 64 entries. A CEL rule rejects the field unless the schema isGCPAnthropicorAWSAnthropic, and onlyanthropic-betais honored today. v1beta1 only.VersionedAPISchema.name:AWSOpenAIandTypeSafe: Two new schema values.TypeSafedefaults its version tov1.GatewayConfig.spec.extProc.metadataForwardingNamespaces: Optional list (max 32) of dynamic metadata namespaces that Envoy forwards to the external processor. It is required forcredentialOverride.fromDynamicMetadata. Under Envoy GatewaymergeGateways, every Gateway in the class shares the union of these namespaces.MCPRoute.spec.prefixModeandbackendRefs[].prefixMode:AlwaysorNever; unset meansAlways. The per-backend value takes precedence. v1beta1 only.MCPRoute.spec.backendRefs[].promptSelector: Filters prompts withinclude,includeRegex,exclude, orexcludeRegex(max 32 each). Exclude rules win over include rules. v1beta1 only.MCPRoute.spec.securityPolicy.mergeTypeandspec.backendTrafficPolicy.mergeType: Envoy GatewayMergeType;Replaceis rejected. Unset keeps the previous override behavior.MCPRoute.spec.backendRefs[].securityPolicy.apiKey.injectionPolicy:Always(default) orIfNotPresent.IfNotPresentcannot be combined withqueryParam. v1beta1 only.MCPRoute.spec.securityPolicy.oauth.authorizationServerMetadataUrl: Optional URI, up to 1024 characters.- New
MCPRoutevalidation for JWT-based CEL:securityPolicy.oauthis now required when an authorization rule orbackendSelectorrule'scelexpression referencesauth.jwt. See Breaking Changes.
Deprecations
cache_creation_input_tokensin translated OpenAI-format responses: Responses translated from Claude and Bedrock now reportcache_write_tokensalongside the legacycache_creation_input_tokensfield. Readcache_write_tokensinstead; the legacy field will be removed in v1.3.0.
⚠️ Breaking Changes
- Random MCP session encryption seed: The Helm chart no longer defaults to the well-known
default-insecure-seed. It generates a random seed, stores it in the Secret<controller fullname>-mcp-session-encryption, and reuses it on later upgrades. Upgrading therefore ends active MCP sessions, and clients must re-initialize. To keep sessions during the upgrade, setcontroller.mcp.sessionEncryption.fallback.seed=default-insecure-seedand remove it once clients have reconnected. If you render the chart withhelm template(common with GitOps), setcontroller.mcp.sessionEncryption.existingSecretorseed. Otherwise every render generates a new seed. - MCP OAuth tokens must match
oauth.issuer: The generated JWT provider now checks the token'sissclaim. Tokens with a missing or different issuer, including a trailing-slash mismatch, get 401. Make suresecurityPolicy.oauth.issuermatches your identity provider'sissexactly. - JWT data in MCP CEL requires
securityPolicy.oauth: The gateway readsrequest.auth.jwtclaims and scopes, and the session subject, only from a JWT that Envoy has verified. WithoutsecurityPolicy.oauth, claims are empty, and new CRD validation rejects authorization orbackendSelectorCEL that referencesauth.jwt. Existing routes that break this rule keep their stored spec, but every update is rejected. At runtime their claim lookups fail, and therefore deny, and their scope checks evaluate to false. Addoauthor remove the JWT references. MCP sessions are now bound to the verified subject; reusing a session ID as a different user returns 401. - MCP CEL evaluation errors now deny: An authorization or
backendSelectorCEL expression that fails at runtime, for example by reading a header that isn't present, used to be skipped. It now denies:tools/callreturns 403, the tool is hidden fromtools/list, and the backend is left out of the session. Write expressions defensively, for example"x-team" in request.headers && request.headers["x-team"] == "blocked". - MCP OAuth discovery honors
spec.headers: On an MCPRoute with bothheadersmatches andoauth, the OAuth well-known endpoints now apply the same header matches. Clients that run OAuth discovery without those headers get 404. - Cross-namespace Secrets in
BackendSecurityPolicyneed a ReferenceGrant:azureCredentials.clientSecretRef,gcpCredentials.credentialsFile.secretRef, and the OIDCclientSecretunder AWS, Azure, and GCPoidcExchangeTokenwere read from other namespaces without any check. ThesecretRefofapiKey,azureAPIKey,anthropicAPIKey, andawsCredentials.credentialsFileignored its namespace, so a cross-namespace reference failed or silently used a Secret with the same name from the policy's own namespace. All of these now honor the namespace, and a missing grant sets the policy to NotAccepted. Already-issued tokens keep working until they expire, so the failure may appear late. Create a ReferenceGrant in the Secret's namespace (see Upgrade Guidance). - Ungranted cross-namespace backends no longer reach the data plane: An
AIGatewayRoutebackendRefto anAIServiceBackendorInferencePoolin another namespace without a ReferenceGrant already marked the route NotAccepted, but its credentials were still written to the external processor's config. Such backends are now left out. Working setups already have the grant; this affects only routes that were already NotAccepted. - ReferenceGrant
to.nameis now enforced: A ReferenceGranttoentry that setsnameused to authorize every object of that group and kind in its namespace. It now authorizes only the named object. References to other objects, such as another Secret,AIServiceBackend, orInferencePool, are NotAccepted and left out of the data plane. Check grants that setnamebefore upgrading. - Per-tenant token limits are now enforced: With the
QuotaPolicyDistinctfix, tenants can start receiving 429 at the limits you configured. Review per-tenant limits before upgrading. - External processor
-configPathremoved: The external processor reads only its config bundle (-configBundlePath), and the controller no longer writes the legacyfilter-config.yamlSecret. Existing legacy Secrets are left in place and can be deleted. This matters only if you run the external processor yourself, or still have Envoy pods created before config bundles existed (v1.0.x era). Restart those pods before upgrading.
🐛 Bug Fixes
- Credential rotation no longer serves stale tokens: Fixes a v1.1.0 regression: the controller read backend credential Secrets from its informer cache, so a rotated credential could miss the filter config and upstream calls failed until the next reconcile. Secrets are read from the API server again. Credential changes in a
BackendSecurityPolicythat targets anInferencePoolnow also reach the Gateway's filter config. - Revoking a ReferenceGrant takes effect immediately: Deleting a ReferenceGrant, or removing a
fromentry, now reconciles the affectedAIGatewayRoutes (includingInferencePoolbackends) andBackendSecurityPolicys (including OIDC client secrets). Access they relied on used to stay in place until an unrelated reconcile. - Gateways with TCP or TLS listeners translate again: A listener without an HTTP connection manager, such as a TCPRoute or TLS-passthrough listener, made the extension server fail translation for the whole Gateway. Those listeners are now skipped.
- Controller crash on combined priority backends: Adding a
backendRefbefore itsAIServiceBackendexisted could panic and crash-loop the controller. The controller now logs the mismatch and continues. - Self-signed webhook survives
helm upgrade: With the default self-signed webhook certificate, an upgrade or Argo CD sync wiped the webhook'scaBundle, so Envoy pods failed admission until the controller restarted. The chart now embeds the CA bundle. credentialOverride.fromDynamicMetadatanow receives metadata: In v1.1.0 Envoy never forwarded the metadata namespace to the external processor. Requests silently used the configured credential, or failed with 401 whenfallbackToConfigured: false. List the namespace inGatewayConfig.spec.extProc.metadataForwardingNamespaces(see Upgrade Guidance).- Token usage is always captured on streams: Duplicate
stream_optionsin a streaming request could stopinclude_usagefrom reaching the provider, which bypassed token accounting. Usage metadata is now also recorded when a client disconnects right after the final SSE frame. Responses API streams framed with CRLF now report usage, model, and tracing data. - Clear 422 errors instead of 500s: Calling an endpoint that a backend's schema doesn't support (for example
/v1/responsesto Vertex AI Gemini) returns 422 Unprocessable Entity. Oversized JSON schemas sent to Vertex AI return 422 once they exceed 50,000 nodes. - Lower external processor memory: Model-name strings used as metric labels no longer keep entire request and response bodies in memory, which was most noticeable with large-context requests.
- Per-tenant token limits charge the right tenant: With
QuotaPolicyDistinctheader selectors, the end-of-stream token cost went into one shared bucket instead of the caller's own. See Breaking Changes for the effect on tenants. - Claude translation fixes: Structured outputs keep the property order you declared. Streamed extended-thinking text is no longer empty. The
web_search_20260209tool on/v1/chat/completionsis now forwarded to Claude instead of rejected with 422; use streaming, because non-streaming responses currently return only the first text block. Tracing of long Claude streams no longer allocates quadratically or panics on malformed delta indices. - MCP interoperability: Backends that send
application/json; charset=utf-8(for example the MCP Java SDK) no longer have their list results dropped. OAuth metadata discovery now treats an empty or non-JSON document as a miss and tries the next well-known URL. aigw runbase URL handling: A path inOPENAI_BASE_URLorANTHROPIC_BASE_URL(for example OpenRouter's/api/v1) is now used as the request prefix instead of being dropped. Base URLs that aren't http or https are rejected.- Unresolvable
InferencePoolroutes are rejected: When Envoy Gateway couldn't resolve anInferencePool, the extension server used to accept a route that only returned 500. It now rejects that route so the error is visible.
📖 Upgrade Guidance
Upgrading from v1.1 needs no CRD storage migration, and every new field is optional. Several security fixes change behavior, though, so work through the steps that apply to you before you upgrade. Upgrade from v1.1.x, not directly from v1.0.x.
1. Upgrade Gateway API and Envoy Gateway first
Envoy Gateway v1.9 requires the Gateway API v1.6 CRDs. If you use the standard channel, move any TCPRoute and UDPRoute manifests to gateway.networking.k8s.io/v1 before you upgrade the CRDs. Then upgrade the CRDs and Envoy Gateway. The Agent Router values file is unchanged:
helm template eg-crds oci://docker.io/envoyproxy/gateway-crds-helm \
--version v1.9.2 \
--set crds.gatewayAPI.enabled=true \
--set crds.envoyGateway.enabled=true \
| kubectl apply --force-conflicts --server-side -f -
helm upgrade -i eg oci://docker.io/envoyproxy/gateway-helm \
--version v1.9.2 \
--namespace envoy-gateway-system \
-f https://raw.githubusercontent.com/theagentrouter/agent-router/v1.2.0/manifests/envoy-gateway-values.yaml
Envoy Gateway v1.9 has its own breaking changes. Read its v1.9.0 and v1.9.2 release notes. The ones most likely to affect AI traffic:
- Lua
EnvoyExtensionPolicyis disabled by default. Enable it withextensionApis.enableLua. mergeTypeon aSecurityPolicyorBackendTrafficPolicyis only accepted on route targets.- Tracing client sampling now defaults to 0%.
- EndpointSlice indexing can increase controller memory.
Do not enable Envoy Gateway's new EnvoyProxy mergeBackends option with Agent Router yet. It shares clusters between routes, which breaks how Agent Router resolves the backend for a request.
2. Plan for the MCP session seed change
The chart now generates a random MCP session encryption seed, so upgrading ends active MCP sessions. Pick one:
-
Accept a reconnect. Do nothing. Clients re-initialize their sessions.
-
Keep sessions during the upgrade. Add the old seed as a fallback in your values for this upgrade, then remove it once clients have reconnected:
controller:mcp:sessionEncryption:fallback:seed: default-insecure-seed -
GitOps or
helm template. The chart can't read back a seed it generated earlier, so it generates a new one on every render. Supply your own Secret instead:kubectl create secret generic mcp-session-seed -n envoy-ai-gateway-system \--from-literal=seed="$(openssl rand -hex 32)"Then set
controller.mcp.sessionEncryption.existingSecret: mcp-session-seed.
3. Check your MCPRoutes
securityPolicy.oauth.issuermust match your identity provider'sissclaim exactly, including any trailing slash.- Any authorization or
backendSelectorCEL that referencesrequest.auth.jwtneedssecurityPolicy.oauthon the same route. - Make CEL expressions null-safe, because evaluation errors now deny. For example, use
"x-team" in request.headers && request.headers["x-team"] == "blocked"instead ofrequest.headers["x-team"] == "blocked". - On routes with both
headersandoauth, make sure clients send those headers during OAuth discovery.
4. Add ReferenceGrants for cross-namespace credential Secrets
If a BackendSecurityPolicy references a credential Secret in another namespace, create a grant in the Secret's namespace. This now applies to every credential type: API keys, AWS credentials files, Azure client secrets, GCP credentials files, and OIDC client secrets.
apiVersion: gateway.networking.k8s.io/v1beta1
kind: ReferenceGrant
metadata:
name: allow-backend-security-policy-secrets
namespace: credentials # the Secret's namespace
spec:
from:
- group: aigateway.envoyproxy.io
kind: BackendSecurityPolicy
namespace: ai-backends # the BackendSecurityPolicy's namespace
to:
- group: ""
kind: Secret
name: openai-api-key # optional: limit the grant to one Secret
A to entry with name now authorizes only that object. Without name, it covers every object of that kind in the namespace. If an existing grant sets name, make sure it lists every Secret, AIServiceBackend, or InferencePool that other namespaces reference.
5. Turn on metadata forwarding for credentialOverride.fromDynamicMetadata
If you use per-request credentials from dynamic metadata, list the metadata namespace in a GatewayConfig, and reference that GatewayConfig from the Gateway with the aigateway.envoyproxy.io/gateway-config annotation:
apiVersion: aigateway.envoyproxy.io/v1beta1
kind: GatewayConfig
metadata:
name: my-gateway-config
namespace: default # same namespace as the Gateway
spec:
extProc:
metadataForwardingNamespaces:
- envoy.filters.http.ext_authz
6. Upgrade Agent Router
helm upgrade -i aieg-crd oci://docker.io/envoyproxy/ai-gateway-crds-helm \
--version v1.2.0 \
--namespace envoy-ai-gateway-system
helm upgrade -i aieg oci://docker.io/envoyproxy/ai-gateway-helm \
--version v1.2.0 \
--namespace envoy-ai-gateway-system
The chart and image names keep the ai-gateway prefix. Only the project name changed.
7. After the upgrade
- Dashboards: GenAI metrics now carry a
gen_ai_backendlabel. Panels that don't aggregate withsum by (...)will split into one series per backend. - Legacy Secrets: the controller no longer writes the old
filter-config.yamlSecrets. You can delete them. - Inference Extension: if you use
InferencePool, upgrade the Gateway API Inference Extension CRDs to v1.6.x and redeploy your endpoint picker.
📦 Dependencies Versions
Updated from Go 1.26.4.
Built on Envoy Gateway v1.9.2, up from v1.8.1.
Envoy Proxy v1.39.1, as shipped by Envoy Gateway v1.9.2.
Updated from v1.5.1. Envoy Gateway v1.9 requires the v1.6 CRDs.
Updated from v1.0.2.
Unchanged from v1.1.
⏩ Patch Releases
🙏 Acknowledgements
v1.2 came from 39 contributors, many of them contributing for the first time. Special thanks to:
- Ignasi Barrera, who led the security hardening across MCP authentication, session binding, and cross-namespace references.
- Hritik Raj, who built the groundwork for the 2026-07-28 MCP specification.
- Deepika Agrawal, Aishwarya Raimule, and Huabing (Robin) Zhao, who added MCP prefix mode, credential injection, policy merging, and OAuth metadata URLs.
- Yang Haoran, hustxiayang, and Xiaolin Lin, who added provider coverage and fixed streaming bugs.
- zirain, who handled the Gateway API Inference Extension v1.6 upgrade and conformance.
- Everyone who reported bugs, reviewed PRs, and joined the community meetings during the move to Agent Router.
🔮 What's Next
Work already in flight for upcoming releases:
- The stateless 2026-07-28 MCP specification, switched on end to end.
- OAuth token exchange for MCP backends, so the gateway can swap a client token for a backend-scoped one.
- A dedicated
MCPBackendCRD, decoupling backend configuration fromMCPRoute. - Longer quota windows (weekly, monthly, and yearly) and quota-aware routing.
- Agent-to-agent (A2A) traffic, building on the A2A preview guide.
The roadmap is community-driven. Join us and help shape it.