> ## Documentation Index
> Fetch the complete documentation index at: https://docs.getbifrost.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# v2.2.0

> v2.2.0 changelog - 2026-09-15

<Tabs>
  <Tab title="NPX">
    ```bash theme={null}
    npx -y @maximhq/bifrost --transport-version v2.2.0
    ```
  </Tab>

  <Tab title="Docker">
    ```bash theme={null}
    docker pull maximhq/bifrost:v2.2.0
    docker run -p 8080:8080 maximhq/bifrost:v2.2.0
    ```
  </Tab>
</Tabs>

<Update label="Bifrost(HTTP)" description="2.2.0">
  ## ✨ Features

  * **Claude Desktop and Cowork Marketplace** - Skills stored in Bifrost can now be registered as a marketplace in Claude Desktop and Cowork, which reject the direct JSON URL and require a cloneable Git repository URL. A new `/api/skills/serve/claude-code.git` endpoint implements the two Git smart-HTTP requests used during a clone and serves a repository containing `.claude-plugin/marketplace.json`; the existing Claude Code flow is unchanged (#7152)
  * **Time-of-Day Peak and Off-Peak Pricing** - Model pricing accepts `off_peak_cost_multiplier` and a `peak_hours` weekly schedule, so providers like DeepSeek that bill the same model at two rates are costed correctly. Base rates are treated as peak; the multiplier scales usage-based charges outside the declared windows. Flat per-request fees, per-search-query fees and guardrail/MCP `AdditionalCost` are never discounted. Windows use IANA timezones, weekday numbers and half-open `HH:MM` intervals that may wrap past midnight, and both fields are editable from the custom pricing override sheet (#6574, #6575, #6576, #6577, #6578, #6579, #7054)
  * **GA Realtime Transcription** - OpenAI and Azure GA transcription sessions are served over both WebSocket and WebRTC with normal Bifrost authentication, routing, governance, guardrails, logging and transcription-aware pricing. These sessions carry only `intent=transcription` on the connection and deliver the routing model later in `session.update` (or in the initial multipart `/v1/realtime/calls` request for WebRTC), so Bifrost now routes on the nested transcription model while preserving realtime connection and turn semantics (#7089)
  * **Regex Model Allow and Block Lists** - `allowed_models` / `blacklisted_models` on virtual keys and `models` / `blacklisted_models` on provider keys accept `regex:<pattern>` entries next to exact names. Patterns are compiled once as case-insensitive full matches, a pattern that is empty, `*` or does not compile is refused with 400, and list-models never surfaces a pattern as a model. Provider-key create and update now validate `models` the same way as `blacklisted_models`. This supersedes the separate `*_patterns` fields, which were added and then withdrawn before release (#6987, #6988, #6989, #6990, #7031, #7133, #7134)
  * **Governance Entity Names on MCP Tool Logs** - `mcp_tool_logs` gains the same attribution shape the `logs` table has: `user_name`, `team_name`, `customer_name` and `business_unit_name` become real columns instead of transients, the multi-valued `team_ids`/`team_names`, `customer_ids`/`customer_names` and `business_unit_ids`/`business_unit_names` sets are stored as index-aligned JSON arrays, and `budget_ids` and `rate_limit_ids` are recorded. Names are written from the request context at ingestion, with nothing resolved on read, so the dashboard stops rendering raw UUIDs (#7154)
  * **Endpoint-Attributed MCP Inspections** - Inspected MCP tool calls are logged with bounded identity sourced from the gateway rather than payload-supplied headers. A new `MCPObservation` carries device, app key, server label, tool name and decision onto both the pending and final log entry, and the MCP logs view falls back to `app_key` when `app` is absent so endpoint-attributed rows show the right app icon and name (#6959)
  * **Virtual MCP References by Name** - Access profiles and governance projects reference Virtual MCPs through `virtual_mcp_name`, making config files portable across environments; names resolve on startup and a name matching nothing is refused. `mcp_configs` (`{ mcp_client_id, tools_to_execute }`) replaces the `mcp_servers` / `mcp_tool_overrides` include-exclude model with a single allowlist, where `["*"]` grants all tools including future ones and `[]` grants none. `virtual_mcp_id` and the old keys are deprecated, still accepted, and folded into the new shape at load time (#7181)
  * **Normalized `error_type` Metric Label** - `bifrost_error_requests_total` gains an `error_type` label with a closed, prefix-structured vocabulary (`caller_*`, `policy_*`, `provider_*`, `bifrost_*`, `_OTHER`) so a 429 from a governance rate limit is distinguishable from a 429 from an upstream, and a 503 Bifrost shed under queue pressure from an upstream overload. Classification resolves a declared `ExtraFields.ErrorType` first, then Bifrost's own markers, then the status code; it deliberately ignores provider `error.type` strings, which disagree across providers for the same condition (#7141)
  * **Bedrock OpenAI-Compatible Endpoint Routing** - A `use_openai_endpoints` flag on Bedrock keys and aliases routes chat completions and responses through Bedrock's `/openai/v1` surface instead of Converse, for models that support it. It is opt-in by design: Converse carries Bedrock Guardrails, `performanceConfig` and `requestMetadata` that the OpenAI-compatible surface silently ignores, so diverting automatically could stop a guardrail from being enforced with no visible error. The alias value wins over the key, matching `use_anthropic_endpoints` precedence (#7071, #7073)
  * **Anthropic Tool Search on Bedrock Claude** - Anthropic tool search (`tool_search_tool_*`, `defer_loading`) is served on `bedrock/` Claude models by routing those requests to InvokeModel / InvokeModelWithResponseStream, the only Bedrock API AWS allows it on; CountTokens counts such requests with the same InvokeModel body. Server-side tool search also survives the Bedrock-native invoke ingress end to end: the tool is carried as an ingress-only marker so the egress predicate can see it, results are returned as a `server_tool_use` plus `tool_search_tool_result` pair rather than a client `tool_use` (which the API rejects when echoed back), the streaming path emits the same pair, replayed blocks round-trip unchanged, and the Anthropic-native response path carries them too (#6900, #6908, #7162, #7163, #7164, #7165, #7166)
  * **Namespace Tool Support Across Providers** - Responses `namespace` tools are flattened in core for every provider whose wire lacks the type, with nested functions renamed to `<namespace>__<function>` so two namespaces sharing a function name no longer collide into an upstream `Tool names must be unique` 400. Returned `function_call` items map back to the bare name plus namespace, prior-turn calls and `tool_choice` names are re-aliased, and a still-duplicate name is rejected with a clear 400 before reaching the provider. Flattened names honour each wire's documented tool-name limit, overridable per model through `tool_name_max_length`. The names a provider reserves for its own server tools come from the datasheet row `reserved_tool_namespaces`, and Codex's literal `functions` namespace is unwrapped to top-level tools for every provider (#7039, #7082, #7084, #7161)
  * **Trusted Networks for the SSRF Guard** - A `trusted_networks` list of IP/CIDR entries the SSRF guard consults before outbound discovery calls, so a self-hosted IdP on an internal network can be reached by the generic provider's discover-endpoints and discover-claims flows. Declaring the key in `config.json` makes it own the whole list, an explicit empty array clears dashboard-added entries, and omitting it leaves the stored allowlist untouched. Hostnames are refused, since DNS would then decide which requests bypass SSRF protection (#7081)
  * **Prompt Cache Reload Through the Server** - `ReloadPromptCache` moves onto `ServerCallbacks` so enterprise can gossip it across nodes. The prompts plugin's in-memory index was previously rebuilt only in the process that served the write, so on a multi-node deployment a prompt published on node A left node B resolving `x-bf-prompt-id` / `x-bf-prompt-version` against a stale index until restart: an unknown version errored and `latest` served the old content. OSS behaviour is unchanged (#7061)
  * **Guardrail Tool-Call Argument Redaction** - Guardrail redaction covers LLM tool-call arguments (Chat function arguments, Responses function arguments and custom-tool input) across the Anthropic streaming and non-streaming paths, reading and writing `delta.partial_json` on `input_json_delta` events and collecting string-valued paths inside a `tool_use` block's `input` without touching tool names, IDs or definitions. A separate identity-based transformer path lets provider-managed rewrites (Model Armor, Bedrock) land in the correct native JSON field even when the same text appears in several fields, verifying `Original` before patching and the written value after (#6977, #7049)
  * **Regions and Service URLs in Plaintext** - Regions and service URLs (Azure endpoint, Vertex/Bedrock/Bedrock Mantle region, vLLM/Ollama/SGL/Databricks URL, MCP connection string) are public identifiers, not credentials, and were being unconditionally redacted into unreadable values in the UI. A new `SecretVar.RedactedIfSecret()` returns a plain clone for a literal value and still masks anything sourced from an env var or vault reference (#7085)
  * **`wait_for_usage` for Custom Providers** - A `wait_for_usage` flag on `custom_provider_config` tells Bifrost the upstream sends a trailing usage-only frame, so the read loop holds open past `finish_reason` until it arrives instead of synthesizing a zero-usage terminal chunk. Termination stays bounded by the usage chunk, two consecutive post-finish heartbeat comments, EOF, or `stream_idle_timeout_in_seconds` (#7187)
  * **Pinnable Log Search Mode** - The logs search box gains a mode dropdown (Auto, Content, Request ID). Auto-detection treated UUID-shaped input as an ID lookup and everything else as a content scan, which breaks for request IDs that are not UUID-shaped and for UUID-shaped strings that should be searched as content. A pinned mode bypasses all sniffing and re-runs the current input immediately (#7149)
  * **MCP Usage Guide Auth Methods** - The MCP usage guide generates client configs for Virtual key, OAuth and Identity provider authentication instead of requiring a virtual key for every config. Credential resolution is centralized in `buildMCPHeaders()`, so OAuth emits no headers, identity provider emits a `Bearer` placeholder, and virtual key keeps `x-bf-vk` (#7111)
  * **Chart Color System** - Dashboard charts, status badges and components read a structured set of CSS custom properties instead of hard-coded hex values, so colors adapt correctly between light and dark themes. Tokens are grouped as semantic (hues 0 to 70 reserved so no category can look like an error), sequential, ordinal for percentile series, and categorical at matched chroma assigned by rank (#7113)

  ## 🐞 Fixed

  * **Client Disconnect Not Cancelling Requests** - A client that closes its socket while Bifrost is still waiting on core (silent upstream header wait, retry backoff) now cancels the request. `ConvertToBifrostContext` starts a socket watcher that peeks the client connection every 500 ms with `MSG_PEEK` and cancels the context on FIN or RST, so upstream retries stop as soon as nobody is listening; previously fasthttp offered no per-request `Done` and a disconnect was only noticed when an SSE write failed. No-op on non-unix platforms and on in-memory test connections (#7035, #7106)
  * **Silent Upstream Never Timed Out** - `default_request_timeout_in_seconds` now bounds the wait for response headers on streaming requests, and cancelling a request closes the upstream socket. Every fasthttp client runs through a Bifrost-owned `RoundTripper` that applies the client's read/write timeouts and the request context to the request write and the header wait, then lifts the socket deadline once headers are parsed so `stream_idle_timeout_in_seconds` remains the only bound on the body. An upstream that accepts the connection and never answers now fails with 504 `RequestTimedOut` and its fallbacks are used, instead of pinning the provider worker. Unary large-response downloads bound every body read the same way, gzip-encoded bodies are classified for large-response mode by decompressed size, and a mid-body connection drop is again reported as the retryable 502 completion-marker error instead of a generic unexpected EOF (#7034, #7104)
  * **Retry Storm After Client Disconnect** - fasthttp-level stale-connection retries no longer multiply `max_retries`. `contextTransport.RoundTrip` reports a pre-header failure on a freshly dialed socket with `retry=false`, so `StaleConnectionRetryIfErr` only walks past pooled keep-alive sockets the upstream closed while idle; an upstream that closes a fresh connection without answering now costs exactly one attempt instead of up to four. Retry backoff also ends as soon as the request context is cancelled, freeing the worker immediately instead of after up to `retry_backoff_max` (#7035, #7105)
  * **Abandoned Request Billing Coin Flip** - Non-streaming requests whose caller had already disconnected were billed and logged only about half the time. The worker's delivery `select` had two simultaneously ready cases, a send into a cap-1 channel and `ctx.Done()`, and Go picks uniformly among ready cases, so terminal post-hooks were skipped roughly 50% of the time. The worker now checks `req.Context.Err()` before the select and calls `billAbandonedTerminal` deterministically, keeping the 5-second timer guard for a caller that leaves between the check and the send (#6972, #7116)
  * **Stream Never Terminated Without `[DONE]`** - An OpenAI-compatible upstream that omits `[DONE]` and then goes silent after `finish_reason` no longer fails the stream when `stream_idle_timeout_in_seconds` fires. The chat and text completion read loops treat an idle timeout after a terminal signal as a parked upstream, abandon the connection rather than drain it, and synthesize the final chunk with the buffered `finish_reason`; a stall before `finish_reason` still surfaces as the idle-timeout error (#7108, #7115)
  * **Dropped SSE Frames in `raw_response`** - Role-only, finish-only and usage-only frames never entered the chunk-forwarding branch, so their bytes were discarded from the reconstructed `raw_response`, leaving the captured audit trail irreconcilable against a provider invoice since the usage frame carries the token counts Bifrost bills from. A `pendingRawFrames` buffer drains onto the next forwarded chunk or the synthetic terminal chunk, the Responses-over-Chat fallback no longer stamps one upstream frame onto every derived event, and `delta.refusal` and `delta.annotations` are forwarded instead of dropped entirely (#7144, #7184)
  * **Bedrock Mantle Trailing Usage** - Bedrock Mantle chat streaming no longer drops the usage-only chunk that arrives after `finish_reason`; `ProviderSendsDoneMarker` now treats `bedrock_mantle` and the legacy Mantle route under the `bedrock` key as sending `[DONE]`, so streamed usage and cost are recorded (#7065, #7076)
  * **Fallbacks Re-Ran the Primary for Image and Video Edits** - Fallbacks for image edit, image variation and video edit requests now reach the configured fallback provider and model. `prepareFallbackRequest` had no arm for those three types, so the shallow copy kept the primary's sub-request pointer and the attempt was routed back to the primary while routing info, headers, the log row and the `fallback_index` metric label all reported it as a fallback. The helper now verifies the prepared request targets the fallback and skips it with a warning otherwise, so a future request type added without an arm fails loudly (#6966, #7118)
  * **Provider Response Headers Leaked Across Fallbacks** - `clearCtxForFallback` missed `BifrostContextKeyProviderResponseHeaders`. Providers set that key before the status check so error paths can forward it, and when a fallback failed pre-flight (key selection failed, a plugin short-circuited, the queue was retiring) nothing overwrote it, so the client received a response attributed to provider B carrying provider A's `Retry-After` and rate-limit headers (#6973, #7021) (thanks [@Huang-404-Q](https://github.com/Huang-404-Q)!)
  * **Credential-Bearing Response Headers Forwarded** - The provider-response extractors filtered only against a fixed map of 28 exact names, so any credential-named header outside it was re-served to the inference caller in both the HTTP response and `extra_fields.provider_response_headers`. The path now also consults `schemas.IsSensitiveHeader`, which recognizes credential names by substring and suffix and already knew about `cf-access-*` and `x-amzn-oidc-*`; a fixed list cannot enumerate the space when `network_config.extra_headers` exists to carry custom auth headers and some upstreams echo request headers back (#7120, #7121) (thanks [@Atharva-Kanherkar](https://github.com/Atharva-Kanherkar)!)
  * **Nil Dereference on Incomplete Fallback Errors** - A plugin returning a `BifrostError` whose nested `Error` field is nil crashed the request worker. Fallback processing now nil-checks the error and guards access to `Error.Type`, using the nil-safe `GetErrorString()` helper, and continues to the next fallback when allowed (#6967, #7110) (thanks [@Constantine3](https://github.com/Constantine3)!)
  * **Bedrock Duplicate Document Names** - Every untitled document block was given the literal default name `document`, and Converse rejects duplicate document names, so any request with two or more untitled documents failed unconditionally with a `ValidationException`. A per-request document namer now disambiguates with numeric suffixes (`document`, `document-2`, and so on) and suffixes titled documents only on an actual collision, on both the Converse and Responses replay paths (#7003, #7027) (thanks [@Huang-404-Q](https://github.com/Huang-404-Q)!)
  * **Bedrock Text Document Source** - Converse rejects document blocks that use a text-only `DocumentSource` unless citations are explicitly enabled, so plain text formats (`text/plain`, `text/markdown`, `text/csv`, `text/html`) failed with `must set one of the following keys: bytes, s3Location`. All document content, including data URLs, percent-encoded payloads and `file_data`, now ships base64-encoded through `source.bytes` (#7072, #7079)
  * **Bedrock Tool Result Images** - Some Bedrock-hosted models (OpenAI and Grok families) reject image blocks placed directly inside a `toolResult` in Converse, even though they accept images in tool output via Responses. A `hoistToolResultImages` pass moves images out of `toolResult` blocks and re-inserts them after the last tool result in the same message, leaving a placeholder text block so the now image-free result is not rejected for being empty. A datasheet row `supports_converse_tool_result_images` overrides the family-level default (#7150)
  * **gpt-oss Message Mistagging on Mantle** - Claude Code replaying a prior assistant message through `POST /anthropic/v1/messages` at a Bedrock Mantle model was rejected with hundreds of validation errors: the Bedrock-grouped ingress converter tagged user and system input text as `output_text` (only `input_text` is valid on input messages) and omitted the required `status` on replayed assistant messages. Bedrock requests with no explicit `max_tokens` now also populate it from the model's known capacity instead of truncating silently (#7074, #7075)
  * **Azure Foundry Output Token Cap** - Azure Foundry deployments of Fireworks-hosted models were silently capped at 4096 output tokens on `/openai/v1/responses` because Microsoft routes those models through chat completions internally. The Azure provider now checks the model's datasheet `supported_endpoints` and transparently serves both `Responses` and `ResponsesStream` through `/openai/v1/chat/completions` when `/v1/responses` is absent. Separately, a turn truncated by the output-token cap on any OpenAI-shaped Responses provider now reports `stop_reason: max_tokens` on the Anthropic egress instead of hiding the truncation as `end_turn` (#6782, #7142)
  * **Gemini Inline Image and Audio Dropped** - Gemini image-generation output (`inlineData`) was silently dropped on `/v1/chat/completions`, both unary and streaming (#7032, #7033) (thanks [@Atharva-Kanherkar](https://github.com/Atharva-Kanherkar)!)
  * **Gemini Image Edit Misclassified** - `isImageEditRequest` only checked `contents[0].parts[0]`, so a request with the prompt text before the `inlineData` part, which is the ordering in Google's own REST edit sample, was misclassified as image generation: the image never reached Vertex and the model invented a picture from the prompt alone, returning HTTP 200. Detection now scans all parts across all contents. `imageConfig.aspectRatio` is also preserved as a typed parameter instead of being folded into a `WxH` size string that collapsed any unsupported ratio to `1:1` (#7173)
  * **Gemini Per-Part Media Resolution Dropped** - `Part` was missing `mediaResolution`, which overrides `generationConfig.mediaResolution` for a single part, and because `Part.UnmarshalJSON` decodes into a closed alias the key was discarded before any conversion ran. Per-part image and PDF tokenization fell back to the model default, so an `ULTRA_HIGH` image billed about 21k prompt tokens through `/genai` instead of about 22.1k direct. The field now round-trips both spellings end to end and is stripped on the OpenAI wire path where it is unknown (#7156)
  * **Gemini `generationConfig` Lost Across Retries** - `convertParamsToGenerationConfigResponses` deleted `top_k`, `frequency_penalty`, `presence_penalty`, `stop_sequences` and `media_resolution` from `ExtraParams` while mapping them into `generationConfig`, and that conversion runs once per attempt on the same request, so every retry or fallback after the first was sent without them. On Vertex the only visible signal was `promptTokenCount` halving on retried requests, making the downgrade silent (#7138) (thanks [@VictorRequenaMaisa](https://github.com/VictorRequenaMaisa)!)
  * **MCP Tool Schema Property Order** - MCP tool schemas keep one property and `$defs` order across tool syncs, so prompt caching is not invalidated by reordering alone (#7169, #7170) (thanks [@dougcalobrisi](https://github.com/dougcalobrisi)!)
  * **vLLM Alias Resolution During Key Selection** - The generic allow and block lists correctly evaluated the user-facing alias, but vLLM key selection compared each key's physical `model_name` against the unresolved alias and rejected valid keys. The alias is now resolved per key before that comparison, so the same public alias can map to different physical model IDs across vLLM instances, while allow and block checks keep matching the original alias. Allowed Models, Blocked Models and Deployments/Aliases are now exposed on the vLLM key form (#6956) (thanks [@Constantine3](https://github.com/Constantine3)!)
  * **opencode-zen Responses Routing** - opencode-zen Responses calls are routed through `/v1/chat/completions`, which its upstream serves, instead of `/v1/responses`, which it does not (#6778, #6819) (thanks [@miguelchico](https://github.com/miguelchico)!)
  * **DeepSeek `max_completion_tokens` Ignored** - DeepSeek's chat-completions endpoint only recognizes the legacy `max_tokens` field and silently ignores `max_completion_tokens`, so the limit had no effect. The field is now remapped on the wire, matching the behaviour already in place for Opencode and Ollama (#7131)
  * **Anthropic Server Tools on Third-Party Endpoints** - Fireworks' Anthropic-compatible endpoint returns 400 when a request includes Anthropic server tools such as `web_search_20250305`, because those run on Anthropic-operated infrastructure that does not exist on third-party hosts; clients whose built-in web search is always on hit this on every request. Unsupported server tools are now dropped before the request leaves Bifrost, the caller's function tools are kept, and the drops are reported on `DroppedUnsupportedTools` instead of failing the call. The same applies to vLLM and SGLang (#7090)
  * **Bedrock Guardrail Headers** - Bedrock's OpenAI-compatible endpoints apply guardrails through request headers rather than the `guardrailConfig` body field Converse uses, so a configured guardrail was silently ignored there. `guardrailIdentifier`, `guardrailVersion` and `trace` are now mapped to the `X-Amzn-Bedrock-Guardrail*` headers on both chat completions and responses, streaming and non-streaming, and the key is consumed so it is not also emitted into the body. A half-formed config with only one of identifier or version is left untouched rather than sent (#7095)
  * **Responses SSE `item: null`** - Responses stream events that carry no item payload no longer serialize `"item": null`, which strict OpenAI Responses clients reject as an invalid frame, breaking streamed `/v1/responses` usage (#6395) (thanks [@ReStranger](https://github.com/ReStranger)!)
  * **Responses `action` String Decode** - `image_generation_call` items where OpenAI emits `action` as a bare JSON string failed to decode, because `UnmarshalJSON` immediately peeked at a `.type` field that cannot be read from a string. That silently dropped the `response.output_item.done` and `response.completed` events carrying the image, leaving the stream without a terminal event and surfacing as a bogus "provider closed the stream" truncation error. The completed item also keeps the generation settings OpenAI echoes back (#7060)
  * **Mid-Conversation System Messages Broke Prompt Caching** - Mid-conversation `role: "system"` messages are inlined in place as `<system-reminder>` user turns on every converter with a top-level system field: Bedrock Converse (Responses and Chat Completions), Gemini Chat Completions, and the Anthropic wire shape used by DeepSeek, Fireworks and SGL. Previously only Claude on Bedrock and Anthropic inlined; everything else hoisted each reminder into the top-level system block. Claude Code appends a trailing `<total_tokens>` reminder after every turn, so the hoisted block grew the front of the prompt each turn and prefix-based caches reported a full cache write and zero cache reads on every turn (#7145)
  * **Unsupported `reasoning.context` Rejected the Request** - A `reasoning.context` value the target model does not accept is dropped on the OpenAI and Azure Responses path, so `all_turns` on the original gpt-5 family including gpt-5-pro, gpt-5.1 to gpt-5.3 and the o-series runs under the model's own `current_turn` default instead of failing with `Unsupported value`. gpt-5.4, gpt-5.5 and gpt-5.6 keep it. Accepted values come from the datasheet row `supported_reasoning_contexts` (#7140)
  * **ClickHouse Retention Filled Replica Disks** - ClickHouse log store deletes no longer run as heavyweight `ALTER TABLE ... DELETE` mutations. The retention cleaner issued one per 100 rows, each rewriting the whole current-month part, and the once-a-minute stale-`processing` sweeps issued one per table unconditionally. Every delete is now a single lightweight `DELETE FROM ... WHERE` per run, skipped when nothing matches. The table TTL derived from `logs_store.retention_days` is reconciled on every startup with a metadata-only `MODIFY TTL`, so changing the value reaches existing tables; `0` leaves an existing TTL untouched (#7098, #7103)
  * **Governance Cleanup Dump Race** - `UsageTracker.Cleanup()` took its final budget and rate-limit snapshots before stopping the periodic reset worker, so `trackerCancel()` could cancel an in-flight dump and fail with `context canceled`. The worker is now cancelled and awaited before the final dumps, a queued ticker event cannot start another reset cycle during shutdown, and `context.Canceled` is treated as expected only when the tracker context was actually cancelled (#7099, #7100) (thanks [@Constantine3](https://github.com/Constantine3)!)
  * **OAuth Refresh Failed for Public Clients** - Public OAuth2 clients registered against servers that only support `token_endpoint_auth_method: none` have no client secret, and unconditionally setting `client_secret=` in the refresh POST body sent an empty `client_secret_post` attempt. Strict authorization servers answered `invalid_client`, flipping the token row to `needs_reauth` even though the refresh token was valid. The parameter is now omitted when the secret is empty, matching the PKCE code-exchange path (#7042)
  * **Complexity Router Skipped Continuation Turns** - When a turn is a continuation, such as a tool result following a prior user message, the complexity router discarded the extracted input and skipped classification entirely if no active session was found, leaving new or recovered sessions without a tier. Continuation turns now keep the populated `ComplexityInput` and fall back to classifying the recovered `LastUserText`; the skip path applies only when that is also empty (#7122)
  * **Runtime Responses-Compat Routing** - Bedrock runtime models that serve the Responses API are routed to it through a dedicated surface resolver rather than falling back to Converse (#7071)
  * **Bedrock Mantle Base Path** - Bedrock Mantle serves each model on exactly one of two base paths (`v1` or `openai/v1`) and returns a 400 on the other. Hard-coded string matching in two packages covered only generations up to GPT-5 and Gemma 4, so GPT-6 and any future closed-generation model silently fell through to the wrong path. Resolution is centralized in `ResolveBedrockMantleBasePath`, backed by a `bedrock_mantle_base_path` datasheet field with family-name detection as the fallback, so a new generation needs a datasheet row rather than a code change (#7077)
  * **Virtual Key `allowed_models: ["*"]` Handling Reverted** - The wildcard handling for governance virtual keys added in #6767 is reverted. Configurations relying on `allowed_models: ["*"]` with an empty synced catalog return to the previous behaviour (#7053)
  * **Helm `perUserHeaderKeys` Not Rendered** - `mcp.clientConfigs[].perUserHeaderKeys` is mapped into the rendered `config.json` (#6033, #6034) (thanks [@CallumWayve](https://github.com/CallumWayve)!)
  * **Helm Plural Access Profiles** - Governance roles accept `access_profiles` as an array in the Helm and config schemas. Bifrost supports multiple access profiles but the schemas accepted only the deprecated singular `access_profile`. The singular form keeps rendering unchanged, the plural wins when both are present, and an explicitly empty plural list clears existing grants (#7044) (thanks [@CarlosLanderas](https://github.com/CarlosLanderas)!)
  * **Sidebar Title Overflow** - Long sidebar item titles are truncated instead of overflowing (#7069)
  * **MCP Logs App Icon** - App icons in the MCP logs table render at 20x20 and no longer shrink when the column is narrow (#7167)

  ## 🗄️ Database Migrations

  * **add\_use\_openai\_endpoints\_column** - Adds the `use_openai_endpoints` column to the provider keys table for Bedrock OpenAI-compatible endpoint routing. Reversible: the rollback drops the added column. Additive and nullable, so it is safe to run during a rolling upgrade.
  * **add\_time\_of\_day\_pricing\_columns** - Adds `off_peak_cost_multiplier` and `peak_hours` to `governance_model_pricing` for time-of-day pricing. Reversible: the rollback drops both added columns. Additive and nullable, so it is safe to run during a rolling upgrade.
  * **mcp\_tool\_logs\_add\_governance\_snapshots** - Adds twelve governance attribution columns to `mcp_tool_logs`: `user_name`, `team_name`, `customer_name`, `business_unit_name`, the `team_ids`/`team_names`, `customer_ids`/`customer_names` and `business_unit_ids`/`business_unit_names` pairs, plus `budget_ids` and `rate_limit_ids`. Reversible: the rollback drops all twelve in reverse order. Additive and nullable, so it is safe to run during a rolling upgrade. The twelve `ALTER TABLE`s run under a bounded DDL lock wait, so startup does not stall behind a long-running log transaction holding `ACCESS EXCLUSIVE` on a continuously written table.

  ## 🐙 Closed GitHub Issues

  * [#6033](https://github.com/maximhq/bifrost/issues/6033) - Helm chart: mcp.clientConfigs\[].perUserHeaderKeys not rendered into config.json
  * [#6778](https://github.com/maximhq/bifrost/issues/6778) - opencode-zen Anthropic endpoint fails, zen upstream doesn't support /v1/responses
  * [#6782](https://github.com/maximhq/bifrost/issues/6782) - Azure Fireworks/Foundry models capped at 4096 output tokens on Responses + Anthropic ingress (chat completions is not); truncation reported as end\_turn
  * [#6825](https://github.com/maximhq/bifrost/issues/6825) - Bedrock provider silently drops Anthropic compaction (compact\_20260112), capability matrix says supported, but Claude egress is Converse-only
  * [#6966](https://github.com/maximhq/bifrost/issues/6966) - fallbacks silently re-target the primary provider for image edit / variation requests
  * [#6967](https://github.com/maximhq/bifrost/issues/6967) - shouldContinueWithFallbacks nil-derefs BifrostError.Error, crashing the process on a plugin-returned error
  * [#6972](https://github.com/maximhq/bifrost/issues/6972) - abandoned-request billing is a \~50% coin flip when a client disconnects mid-request
  * [#6973](https://github.com/maximhq/bifrost/issues/6973) - provider response headers leak across fallback boundaries (clearCtxForFallback misses ProviderResponseHeaders)
  * [#7003](https://github.com/maximhq/bifrost/issues/7003) - Bedrock Converse assigns duplicate default name "document" to untitled document blocks, ValidationException
  * [#7032](https://github.com/maximhq/bifrost/issues/7032) - Gemini image-generation output (inlineData) silently dropped on /v1/chat/completions, both unary and streaming
  * [#7034](https://github.com/maximhq/bifrost/issues/7034) - `default_request_timeout_in_seconds` and `stream_idle_timeout_in_seconds` do not fire while waiting for response headers, a silent upstream blocks the request until the upstream closes, and the fallback is never used
  * [#7035](https://github.com/maximhq/bifrost/issues/7035) - upstream retries continue after the client has disconnected, and go past `max_retries`, one abandoned request produced 10 upstream attempts over \~20 minutes
  * [#7048](https://github.com/maximhq/bifrost/issues/7048) - Compat namespace flattening creates duplicate tool names for DeepSeek 4.1 Flash
  * [#7065](https://github.com/maximhq/bifrost/issues/7065) - Bedrock Mantle chat streaming drops trailing usage after finish\_reason
  * [#7072](https://github.com/maximhq/bifrost/issues/7072) - Bedrock Converse drops text-format document bytes, DocumentSource "must set one of the following keys: bytes, s3Location" (v1 to v2 regression)
  * [#7074](https://github.com/maximhq/bifrost/issues/7074) - openai.gpt-oss-120b via Bedrock Responses API mistags replayed history as output\_text instead of input\_text, breaks multi-turn Claude Code sessions
  * [#7098](https://github.com/maximhq/bifrost/issues/7098) - ClickHouse logs store: retention cleaner runs one `ALTER TABLE ... DELETE` mutation per 100 rows and fills replica disks
  * [#7099](https://github.com/maximhq/bifrost/issues/7099) - UsageTracker cleanup races the periodic rate-limit dump during shutdown
  * [#7108](https://github.com/maximhq/bifrost/issues/7108) - Custom-provider streaming never terminates when the upstream omits \[DONE] (heartbeats mask stream\_idle\_timeout\_in\_seconds)
  * [#7120](https://github.com/maximhq/bifrost/issues/7120) - Provider response-header filter ignores IsSensitiveHeader, forwarding credential-named headers to inference callers
  * [#7143](https://github.com/maximhq/bifrost/issues/7143) - does\_not\_send\_done\_marker drops trailing Chat Completions usage and records zero cost
  * [#7144](https://github.com/maximhq/bifrost/issues/7144) - Chat Completions streaming raw\_response omits usage-only and finish-only SSE frames
  * [#7155](https://github.com/maximhq/bifrost/issues/7155) - Bedrock-native invoke ingress silently drops Anthropic tool search (`tool_search_tool_*` / `defer_loading`), served eagerly over Converse
  * [#7169](https://github.com/maximhq/bifrost/issues/7169) - MCP tool schema property order changes between tool syncs, breaking prompt caching
</Update>

<Update label="Core" description="1.9.0">
  * \[fix]: MCP tool schemas keep one property order across tool syncs [@dougcalobrisi](https://github.com/dougcalobrisi)
  * feat: `WhiteList` and `BlackList` entries that start with `regex:` are RE2 patterns. `MatchEntry` compiles them once as a case-insensitive full match, `Validate` refuses a pattern that is empty, `*` or does not compile, and the list-models pipeline never surfaces a pattern as a model. Plain entries keep their exact, case-insensitive match
  * fix: fasthttp-level stale-connection retries no longer multiply `max_retries`. `contextTransport.RoundTrip` reports a failure before response headers on a freshly dialed socket with `retry=false`, so `network.StaleConnectionRetryIfErr` only walks past pooled keep-alive sockets the upstream closed while idle; an upstream that closes a fresh connection without answering now costs exactly one attempt of Bifrost's own retry budget instead of up to four. The retry backoff also ends as soon as the request context is cancelled, so a worker is freed immediately instead of after up to `retry_backoff_max` (#7035)
  * \[fix]: Gemini provider - preserve inline image and audio data in chat completions responses [@Atharva-Kanherkar](https://github.com/Atharva-Kanherkar)
  * \[fix]: route opencode-zen Responses calls through /v1/chat/completions, which its upstream does not serve on /v1/responses [@miguelchico](https://github.com/miguelchico)
  * feat: the namespace-tool names a provider reserves for its own server tools (`web`, `image_gen`, `browser`, `python` on Bedrock and Bedrock Mantle) are now read from the datasheet row `reserved_tool_namespaces` (new `ModelCapabilities` field and `ModelCaps.ReservedToolNamespaces`) for the base provider and canonical model; a non-empty row replaces the hardcoded list, an absent or empty row keeps it, so any OpenAI-wire provider can reserve names through a row without a code change
  * fix: Codex >= 0.147 wraps its default tools in a namespace literally named `functions` (openai/codex#37022), which Bedrock Mantle reserves and rejects with `User-defined namespace 'functions' collides with an existing tool namespace`. Core now unwraps that namespace to top-level tools, unprefixed, for every provider before dispatch; Codex treats the bare name and the explicit `functions` namespace as the same tool, so no mapping back is needed; a description on the `functions` namespace is prepended to each hoisted member, as flattening does
  * fix: Responses `namespace` tools are now flattened in core for every provider whose wire does not support the type (Anthropic, Gemini, Vertex, Bedrock Converse, DeepSeek, and every OpenAI-compatible third party), with nested functions renamed to `<namespace>__<function>` so two namespaces that share a function name no longer collide into an upstream `Tool names must be unique` 400; returned `function_call` items are mapped back to the bare `name` plus `namespace` (unary and streaming), prior-turn calls carrying `namespace` and `tool_choice` names are re-aliased to match, and a name that is still duplicated after flattening or a `tool_choice` that matches several namespaces is rejected with a clear 400 before reaching the provider. Bedrock answers namespace support from its own surface resolver via the new optional `ResponsesNamespaceToolProvider` interface. A datasheet row `supports_namespace_tools` (new `ModelCapabilities` field and `ModelCaps.SupportsNamespaceTools`) overrides the per-provider default for a (provider, model) pair; with no row the default applies. A row can only narrow within what the wire can carry: the Anthropic Messages API (Anthropic, Claude on Azure or Bedrock Mantle), the Gemini API (Gemini, Vertex) and Bedrock Converse have no namespace container, so they answer false regardless of the row and always flatten. Flattened names honour the target wire's documented tool-name limit, 64 characters of `[A-Za-z0-9_-]` for OpenAI-compatible wires, Bedrock Converse and Fireworks, 128 for Anthropic, and 128 with `.` and `:` allowed for Gemini and Vertex, overridable per model through the datasheet row `tool_name_max_length` (a row below 10 cannot hold the hashed form and is ignored); a longer alias becomes an 8-hex hash prefix plus the function name, deterministically, so history and `tool_choice` re-alias to the same string. The alias map travels on the prepared request (`BifrostResponsesRequest.NamespaceToolAliases`) and is applied to that attempt's response, unary and streaming; nothing is kept on the request context or in process-wide state (#7048)
  * fix: drop namespace tools whose name Amazon Bedrock reserves (`web`, `image_gen`, `browser`, `python`) on the Bedrock and Bedrock Mantle Responses paths instead of forwarding them into a `tools.namespace` collision 400; a dropped Codex `web` namespace becomes the hosted `web_search` tool on Bedrock Mantle
  * fix: Bedrock Mantle chat streaming no longer drops the usage-only chunk that arrives after `finish_reason`; `ProviderSendsDoneMarker` now treats `bedrock_mantle` (and the legacy Mantle route under the `bedrock` key) as sending `[DONE]`, so streamed usage and cost are recorded (#7065)
  * fix: `default_request_timeout_in_seconds` now bounds the wait for response headers on streaming requests, and cancelling a request now closes the upstream socket. Every fasthttp client is driven through a Bifrost-owned `RoundTripper` (`core/providers/utils/roundtripper.go`) that applies the client's read/write timeouts and the request context to the request write and the header wait, then lifts the socket deadline once headers are parsed so `stream_idle_timeout_in_seconds` remains the only bound on the body. `BuildStreamingClient` and `BuildLargeResponseClient` no longer zero the timeouts. An upstream that accepts the connection and never answers now fails with 504 `RequestTimedOut` after the request timeout and the request's fallbacks are used, instead of pinning the provider worker until the upstream closes the socket; `MakeRequestWithContext` no longer leaves a background `client.Do` running until `ReadTimeout` after its context is cancelled (#7034)
  * fix: unary large-response downloads (`FinalizeResponseWithLargeDetection`) now bound every body read with `stream_idle_timeout_in_seconds` and close the upstream socket on request cancellation, mirroring the streamed passthrough path; a stall during the prefetch surfaces as 504 `RequestTimedOut` instead of pinning the provider worker, and a stall while the transport drains the `LargeResponseReader` fails the read with the idle-timeout error instead of blocking the writer indefinitely (#7104)
  * fix: gzip-encoded unary responses are now classified for large-response mode by decompressed size; `Content-Length` describes the compressed bytes, so a body that was small on the wire but large once decompressed used to be materialized in full past `large_response_threshold` (#7104)
  * fix: an OpenAI-compatible upstream that omits `[DONE]` and then goes silent after `finish_reason` no longer fails the stream when `stream_idle_timeout_in_seconds` fires. The chat and text completion read loops now treat an idle timeout after a terminal signal as a parked upstream, mark the stream parked so the connection is abandoned rather than drained, and synthesize the final chunk with the buffered `finish_reason`; a stall before `finish_reason` still surfaces as the idle-timeout error (#7108)
  * fix: a streamed upstream that drops the connection mid-body is again reported as the retryable 502 `provider closed the stream before sending a completion marker` error instead of a generic `Error reading stream: unexpected EOF`. The Bifrost round tripper's chunked decoder surfaced the drop as `io.ErrUnexpectedEOF`, which no provider read loop treats as end of stream; it now reports the plain `io.EOF` fasthttp always did and discards the half-read connection (#7104 follow-up)
  * fix: fallbacks for image edit, image variation and video edit requests now reach the configured fallback provider and model. `prepareFallbackRequest` had no arm for those three types, so the shallow request copy kept the primary's sub-request pointer and the "fallback" attempt was routed back to the primary while `RoutingInfo`, the `x-bifrost-routing-info-*` headers, the log row and the `fallback_index` metric label reported it as a fallback. The helper now also verifies the prepared request targets the fallback provider and model and skips the fallback with a warning otherwise, so a future request type added without an arm fails loudly instead of silently re-running the primary (#6966)
  * fix: route Bedrock Claude requests that carry a `compact_20260112` edit to InvokeModel / InvokeModelWithResponseStream with the native Anthropic Messages body, so server-side compaction works on `bedrock/` models, including keys that pin an inference-profile ARN. AWS documents compaction as unsupported on the Converse API, which previously received the edit and silently ignored it (#6825)
  * feat: serve Anthropic tool search (`tool_search_tool_*`, `defer_loading`) on `bedrock/` Claude models by routing those requests to InvokeModel / InvokeModelWithResponseStream, the only Bedrock API AWS allows it on; CountTokens counts such requests with the same InvokeModel body via the `invokeModel` input
  * fix: a `reasoning.context` value the target model does not accept is now dropped on the OpenAI and Azure Responses path, so `all_turns` on the original gpt-5 family incl. gpt-5-pro, gpt-5.1 to gpt-5.3, and the o-series runs under the model's own `current_turn` default instead of failing with `Unsupported value: 'all_turns' is not supported with the 'gpt-5-pro' model`; gpt-5.4, gpt-5.5 and gpt-5.6 keep it. The accepted values come from the datasheet row `supported_reasoning_contexts` (new `ModelCapabilities` field and `ModelCaps.SupportedReasoningContexts`), falling back to a name-based default of `auto` and `current_turn` everywhere plus `all_turns` on gpt-5.4+
  * fix: mid-conversation `role: "system"` messages are now inlined in place as `<system-reminder>` user turns for every model family on every converter that has a top-level system field: Bedrock Converse (Responses and Chat Completions), Gemini Chat Completions (`systemInstruction`), and the Anthropic wire shape used by DeepSeek, Fireworks and SGL. Previously only Claude on Bedrock/Anthropic inlined; everything else hoisted each reminder into the top-level system block. Claude Code appends a trailing `<total_tokens>` system reminder after every turn, so the hoisted block grew the front of the prompt each turn and prefix-based caches (Bedrock implicit caching for `global.openai.gpt-5.6-luna` on Converse, Gemini implicit caching, DeepSeek context caching) reported a full cache write and zero cache reads on every turn
  * \[fix]: vLLM key selection resolves the requested alias through `key.Aliases` per key before comparing it with `vllm_key_config.model_name`, so a valid key is no longer rejected when the request used an alias; the same public alias can map to different physical model IDs across vLLM instances, and allow and block checks keep matching the original user-facing alias [@Constantine3](https://github.com/Constantine3)
  * \[fix]: Gemini Responses requests keep `top_k`, `frequency_penalty`, `presence_penalty`, `stop_sequences` and `media_resolution` across retries and fallbacks. `convertParamsToGenerationConfigResponses` deleted those keys from `ExtraParams` while mapping them into `generationConfig`, and the conversion runs once per attempt on the same request, so every attempt after the first reached the provider without them; the outbound `ExtraParams` is now built through a filtered copy instead of mutating the request [@VictorRequenaMaisa](https://github.com/VictorRequenaMaisa)
  * feat: `wait_for_usage` on `custom_provider_config` keeps the OpenAI-compatible chat and text completion read loops open past `finish_reason` until the trailing usage-only frame arrives, so a custom provider that also sets `does_not_send_done_marker` no longer synthesizes a zero-usage terminal chunk and records the request at zero tokens and zero cost. `WaitForUsage` is carried on `BifrostContextKeyWaitForUsage` with the same set-or-clear discipline as `DoesNotSendDoneMarker`, so it never leaks onto a fallback provider that did not declare it, and termination stays bounded by the usage chunk, two consecutive post-finish heartbeat comments, EOF, or `stream_idle_timeout_in_seconds` (#7143)
  * fix: `ExtraFields.RawResponse` now carries role-only, finish-only and usage-only SSE frames on the OpenAI-compatible chat stream. Those frames never entered the chunk-forwarding branch, so their bytes were discarded from the reconstructed `raw_response`, which made the captured audit trail irreconcilable against a provider invoice because the usage frame carries the token counts Bifrost bills from. A `pendingRawFrames` buffer accumulates them and drains onto the next forwarded chunk or the synthetic terminal chunk; the Responses-over-Chat fallback no longer stamps one upstream frame onto every derived event, and `delta.refusal` and `delta.annotations` are added to the forwarding predicate so refusals and streamed URL citations are no longer dropped entirely (#7144)
  * fix: Anthropic server-side tool search survives the Bedrock-native invoke ingress end to end. `tool_search_tool_*` was dropped during the Converse-shaped intermediate conversion, leaving the egress predicate blind to the tool and routing the request to Converse where server-side tool search is unavailable; the tool is now carried as an ingress-only marker (`json:"-"`, never reaching a Converse body) and the neutral tool is rebuilt in `ToBifrostResponsesRequest` along with `defer_loading`. Results are returned as a `server_tool_use` plus `tool_search_tool_result` pair rather than a client `tool_use`, which Anthropic rejects when the caller echoes a `tool_result` for the `srvtoolu_` ID on the next turn, on both the unary and streaming paths; `BedrockContentBlock.UnmarshalJSON` gained `server_tool_use` and `tool_search_tool_result` cases so replayed turns round-trip unchanged instead of falling through to empty structs; and the Anthropic-native response path carries the same blocks (#7155)
  * fix: native Gemini `generateContent` image requests routed through `/genai` no longer return HTTP 200 while silently ignoring part of the request. `isImageEditRequest` only inspected `contents[0].parts[0]`, so a request with the prompt text before the `inlineData` part, the ordering in Google's own REST edit sample, was misclassified as image generation: the image never reached Vertex and the model invented a picture from the prompt alone. Detection now scans every part across every content. `imageConfig.aspectRatio` is preserved as a typed `aspect_ratio` param that both outbound converters prefer, instead of being folded into a `WxH` size string that collapsed any ratio outside `1:1`, `3:4`, `4:3`, `9:16` and `16:9` back to square (#7173)
  * fix: Gemini `Part` gained `mediaResolution`, the per-part override for `generationConfig.mediaResolution`. `Part.UnmarshalJSON` decodes into a closed alias, so the key was discarded before any conversion ran and per-part image and PDF tokenization fell back to the model default; an `ULTRA_HIGH` image billed about 21k prompt tokens through `/genai` instead of about 22.1k direct. Both the `mediaResolution` and `media_resolution` spellings are parsed, the value round-trips through `ResponsesMessageContentBlock` and is reconstructed on the outbound part, and it is stripped on the OpenAI wire path where the field is unknown (#7156)
  * fix: image blocks inside a Bedrock Converse `toolResult` are hoisted out and re-inserted immediately after the last tool result in the same message, with a placeholder text block left behind so the now image-free result is not rejected for being empty. Some Bedrock-hosted models, notably the OpenAI and Grok families, reject nested tool-result images on Converse even though they accept images in tool output via the Responses API. `BedrockModelSupportsToolResultImages` supplies the name-based default and the datasheet field `supports_converse_tool_result_images` (new `ModelCapabilities` field and `ModelCaps.SupportsConverseToolResultImages`) overrides it per row (#7150)
  * feat: `bifrost_error_requests_total` gains a normalized `error_type` label drawn from a closed, prefix-structured vocabulary (`caller_*`, `policy_*`, `provider_*`, `bifrost_*`, `_OTHER`), so a 429 from a governance rate limit is distinguishable from a 429 from an upstream provider, a 403 from a blocked model from one from a bad key, and a 503 Bifrost shed under queue pressure from an upstream overload. `ClassifyErrorType` resolves a declared `ExtraFields.ErrorType` first, then Bifrost's own string markers, then the upstream status code; it deliberately ignores the provider's own `error.type` and `error.code` strings, which disagree across providers for the same condition. Values are prefixed by fault domain so an alarm expression can match a family with a single regex (#7141)
  * fix: Azure Foundry deployments of Fireworks-hosted models are no longer silently capped at 4096 output tokens on `/openai/v1/responses`, where Microsoft routes those models through chat completions internally. The Azure provider checks the model's datasheet `supported_endpoints` and, when `/v1/responses` is absent, transparently serves both `Responses` and `ResponsesStream` through `/openai/v1/chat/completions` and converts the result back to the Responses shape; a model with an explicit `/v1/responses` entry or no datasheet row at all is unchanged. Separately, `ToAnthropicResponsesResponse` now follows `StopReason > IncompleteDetails > tool_use inference > end_turn`, so a turn truncated by the output-token cap on any OpenAI-shaped Responses provider reports `stop_reason: max_tokens` instead of hiding the truncation as `end_turn` (#6782)
  * fix: DeepSeek chat requests move `max_completion_tokens` into `max_tokens` on the wire. DeepSeek's chat-completions endpoint only recognizes the legacy field and silently ignores `max_completion_tokens`, so the limit had no effect; this matches the behaviour already in place for Opencode and Ollama (#7131)
  * feat: guardrail redaction gains an identity-based transformation path for provider-native request and response bodies, so exact provider-managed rewrites such as Model Armor and Bedrock land in the correct native JSON field even when the same text value appears in several fields. `TextTargetID`, `TextTargetIDForIndex` and `TextRewrite` identify each guardrail-visible field by a stable ID rather than by content value; `RawRequestBodyTextTransformer` and `RawResponseTextTransformer` let integrations register handlers on the request context; and `rewriteRawJSONTextTargets` validates target uniqueness, verifies each `Original` value before patching and re-reads the written value after, so stale or mismatched normalized text cannot silently rewrite the wrong field (#7049)
  * feat: guardrail redaction covers LLM tool-call arguments, Chat function arguments, Responses function arguments and custom-tool input, across the Anthropic streaming and non-streaming paths. `input_json_delta` events are inspected and rewritten alongside `text_delta`, reading and writing `delta.partial_json`; `collectAnthropicArgumentStringPaths` recursively collects string-valued paths inside a `tool_use` block's `input` without touching tool names, IDs or definitions; `BifrostResponsesStreamResponse` gained `Input` to carry the full custom-tool payload on `custom_tool_call_input.done`; and `DeepCopyResponsesMessage` deep-copies `ResponsesCustomToolCall` so a copied message cannot mutate the original (#6977)
  * fix: an abandoned non-streaming request, one whose caller context was already cancelled by the time the upstream finished, is billed and logged deterministically instead of about half the time. The worker's delivery `select` had two simultaneously ready cases, a send into a cap-1 channel and `ctx.Done()`, and Go picks uniformly among ready cases, so the terminal post-hooks that record billing and finalize the log row were skipped roughly 50% of the time. `requestWorker` now checks `req.Context.Err()` before the select on both the error and success paths and calls `billAbandonedTerminal` directly, keeping the 5-second timer guard inside the select for a caller that leaves between the check and the send (#6972)
  * fix: provider response headers are filtered by classifier, not only by a fixed map of 28 exact names, so a credential-named header outside that map is no longer re-served to the inference caller in the HTTP response and `extra_fields.provider_response_headers`. The extractors now also consult `schemas.IsSensitiveHeader`, which matches credential names by substring and suffix and already recognized `cf-access-*` and `x-amzn-oidc-*`; a fixed list cannot enumerate the space when `network_config.extra_headers` exists to carry custom authentication headers and some upstreams echo request headers back (#7120) [@Atharva-Kanherkar](https://github.com/Atharva-Kanherkar)
  * \[fix]: a plugin that returns an incomplete `BifrostError` whose nested `Error` field is nil no longer panics the request worker. The fallback decision helper nil-checks `fallbackErr`, guards access to `Error.Type` and reads the message through the nil-safe `GetErrorString()` helper, continuing to the next fallback when allowed; behaviour for cancelled requests and `AllowFallbacks == false` is unchanged (#6967) [@Constantine3](https://github.com/Constantine3)
  * \[fix]: `clearCtxForFallback` now clears `BifrostContextKeyProviderResponseHeaders`. Providers set that key from their own HTTP response before the status check so error paths can forward it, and when a fallback attempt failed pre-flight, key selection failing for the fallback provider, a plugin short-circuiting it, or the queue retiring, nothing overwrote the key and the primary's headers survived onto the fallback's error response. A client could receive a response attributed to provider B carrying provider A's `Retry-After` and `x-ratelimit-remaining-*`, and wait according to a limit belonging to a provider that never served the request (#6973) [@Huang-404-Q](https://github.com/Huang-404-Q)
  * \[fix]: untitled document blocks get unique names on the Bedrock Converse path. Every untitled block was named the literal `document`, and Converse rejects duplicate document names, so any request carrying two or more untitled documents failed unconditionally with `ValidationException: Messages can't contain duplicate document names`. A per-request namer threaded through the content-block conversion disambiguates with numeric suffixes (`document`, `document-2`, `document-3`) and suffixes explicitly titled documents only on an actual collision; the Responses input replay path gets the same treatment scoped to its content-block list (#7003) [@Huang-404-Q](https://github.com/Huang-404-Q)
  * \[fix]: `BifrostResponsesStreamResponse.Item` is `omitempty`, so Responses stream events that carry no item payload no longer serialize `"item": null`. Strict OpenAI Responses clients reject those as invalid stream frames, which broke streamed `/v1/responses` usage entirely; `response.output_item.added` and `response.output_item.done` still emit the item object [@ReStranger](https://github.com/ReStranger)
  * feat: `SecretVar.RedactedIfSecret()` returns a plain clone when the value is a literal and delegates to `Redacted()` when it is env-var or vault-backed, so regions and service URLs (Azure endpoint, Vertex/Bedrock/Bedrock Mantle region, vLLM/Ollama/SGL/Databricks URL, MCP connection string) are readable in the UI instead of being masked as if they were credentials. It always returns a fresh pointer, so the redacted copy never aliases the live config, and it replaces the ad-hoc `IsFromSecret()` guard blocks that tried to replicate this inline (#7085)
  * feat: GA realtime transcription is served for OpenAI and Azure over both WebSocket and WebRTC with normal Bifrost authentication, routing, governance, guardrails, logging and transcription-aware pricing. Unlike a normal realtime session, which names its routing model in the connect URL, a GA transcription session carries only `intent=transcription` and delivers the model later in `session.update` under `audio.input.transcription.model` (or in the initial multipart `/v1/realtime/calls` request on WebRTC), so routing now resolves from the nested transcription model while realtime connection and turn semantics are preserved (#7089)
  * fix: Claude Code replaying a prior assistant message through `POST /anthropic/v1/messages` at a Bedrock Mantle model is no longer rejected with hundreds of validation errors. The Bedrock-grouped ingress converter tagged user and system input text as `output_text`, though only `input_text` is valid on input messages, and omitted the required `status` on replayed assistant output messages, so Mantle's strict OpenAI-compatible validator refused the whole request. Replayed assistant messages now carry `Status: "completed"` and input text blocks are tagged `input_text`. Bedrock requests with no explicit `max_tokens` also populate it from the model's known capacity instead of truncating silently on large context windows (#7074)
  * fix: Anthropic server tools such as `web_search_20250305` are stripped before a request reaches Fireworks, vLLM or SGLang on their Anthropic-compatible endpoints, which return 400 for tools executed by Anthropic-operated infrastructure that does not exist on third-party hosts; clients whose built-in web search is always on, such as Codex, hit this on every request. `ProviderFeatures` entries for `Fireworks`, `VLLM` and `SGL` declare all server-tool flags off so the existing validators strip them, `StripUnsupportedServerToolsFromRawBody` mirrors that on the raw-body passthrough path, the caller's function tools are kept, and the drops are reported on the response's `DroppedUnsupportedTools` instead of failing the call (#7090)
  * fix: `guardrailConfig` reaches Bedrock's OpenAI-compatible `chat/completions` and `responses` surfaces, streaming and non-streaming, which apply guardrails through request headers rather than the body field Converse uses, so a configured guardrail was previously ignored there with no error. `withGuardrailHeaders` maps `guardrailIdentifier`, `guardrailVersion` and optional `trace` to `X-Amzn-Bedrock-GuardrailIdentifier`, `X-Amzn-Bedrock-GuardrailVersion` and `X-Amzn-Bedrock-Trace`, then deletes the key from `ExtraParams` so it is not also emitted into the body. A half-formed config with only one of identifier or version is left untouched rather than sent, since both are required upstream, and the base header map is cloned rather than mutated (#7095)
  * feat: `use_openai_endpoints` on Bedrock keys and aliases routes chat completions and responses through Bedrock's OpenAI-compatible `/openai/v1` endpoints instead of Converse, for models that support them, mirroring the existing `use_anthropic_endpoints` pattern. It is opt-in by design: Converse carries Bedrock Guardrails, `performanceConfig` and `requestMetadata` that the OpenAI-compatible surface silently ignores, so diverting automatically could stop a guardrail from being enforced with no visible error. `ResolveUseOpenAIEndpoints` gives the alias value precedence over the key, and the narrow `runtimeServesResponses` check is replaced by a general `runtimeServesOpenAIAPI` that takes a `BedrockAPI` discriminator; chat completions, streaming and non-streaming, gained the runtime path that previously existed only for Responses (#7073)
  * fix: Bedrock runtime models that serve the Responses API are routed to it through a dedicated surface resolver rather than falling back to Converse (#7071)
  * fix: Bedrock Converse document blocks always ship their content base64-encoded through `source.bytes`, never through `source.text`, which Converse rejects unless citations are explicitly enabled. Plain text formats (`text/plain`, `text/markdown`, `text/csv`, `text/html`) previously took the text branch and failed with `DocumentSource must set one of the following keys: bytes, s3Location`; the branch is removed for data URLs, percent-encoded payloads and the non-data-URL `file_data` path alike (#7072)
  * fix: the Bedrock Mantle base path is resolved from the model capabilities datasheet instead of hard-coded string matching in two packages. Mantle serves each model on exactly one of `v1` or `openai/v1` and returns a 400 on the other, and the previous matching covered only generations up to GPT-5 and Gemma 4, so GPT-6 and any future closed-generation model silently fell through to the wrong path. `ResolveBedrockMantleBasePath` applies family-name detection as a fallback and defers to the new `BedrockMantleBasePath` field on `ModelCapabilities`, so a new generation needs a datasheet row rather than a code change (#7077)
  * fix: `image_generation_call` items where OpenAI emits `action` as a bare JSON string, such as `"generate"`, decode correctly. `ResponsesToolMessageActionStruct.UnmarshalJSON` immediately peeked at a `.type` field, which cannot be read from a JSON string, so the whole decode failed with `failed to peek at type field`; that silently dropped the `response.output_item.done` and `response.completed` events carrying the image, leaving the stream without a terminal event and surfacing as a bogus `provider closed the stream` truncation error. The action is now tried as a bare string before falling back to the object type-peek, and `ResponsesImageGenerationCall` keeps the `background`, `output_format`, `quality`, `revised_prompt` and `size` settings OpenAI echoes back on completed items (#7060)
</Update>

<Update label="Framework" description="1.7.0">
  * feat: model allow and block lists accept `regex:` entries. The provider-key aggregate keeps an exact name next to a pattern that also covers it, the catalog allow check tries a pattern against the bare name and `provider/model`, and listings skip pattern entries
  * fix: ClickHouse log store deletes no longer run as heavyweight `ALTER TABLE ... DELETE` mutations. The retention cleaner issued one such mutation per 100 rows, each rewriting the whole current-month part, and the once-a-minute stale-`processing` sweeps issued one per table unconditionally, filling replica disks in minutes. Every delete on the ClickHouse store (retention sweep, `Flush`/`FlushMCPToolLogs`, UI log deletes, async job and webhook delivery expiry) is now a single lightweight `DELETE FROM ... WHERE` per run, skipped entirely when nothing matches. The table TTL derived from `logs_store.retention_days` is now reconciled on every startup with `MODIFY TTL` (metadata only), so changing the value reaches existing tables; `0` leaves an existing TTL untouched (#7098)
  * feat: access profiles and governance projects reference Virtual MCPs through `virtual_mcp_name`, so config files are portable across environments instead of carrying database-assigned integer IDs. Names resolve to stored records on startup and a name that matches nothing is refused; `virtual_mcp_id` is deprecated, still accepted, and wins when both are set. `mcp_configs` (`{ mcp_client_id, tools_to_execute }`) replaces the `mcp_servers` and `mcp_tool_overrides` include-exclude model with a single allowlist, where `["*"]` grants all tools including future ones, `[]` grants none, and a named list grants only those; the old keys are deprecated, still accepted, and folded into the new shape at load time (#7181)
  * feat: `MCPToolLog` records governance entity names beside their IDs, so MCP tool logs carry the same attribution shape the `logs` table has instead of rendering raw UUIDs in the dashboard. `user_name`, `team_name`, `customer_name` and `business_unit_name` stop being `gorm:"-"` transients and become storage; the multi-valued `team_ids`/`team_names`, `customer_ids`/`customer_names` and `business_unit_ids`/`business_unit_names` sets are stored as index-aligned JSON arrays; and `budget_ids` and `rate_limit_ids` are recorded id-only, as in `logs`. Names are written from the request context at ingestion through the new `MCPToolLog.ApplyGovernance` in `framework/logstore/governance.go`, with nothing resolved on read. Added by migration `mcp_tool_logs_add_governance_snapshots` (#7154)
  * feat: endpoint-attributed MCP observations are carried in the standard logging pipeline, so inspected MCP tool calls are logged with bounded identity (device, app key, server label, tool name, decision) sourced from the gateway rather than from payload-supplied headers. `SetMCPObservation` attaches the attribution to a `BifrostContext` and `applyMCPObservation` writes it onto the `MCPToolLog` from both `PreMCPHook` and `PostMCPHook`, snapshotting the observation so it cannot alias across async log entries (#6959)
  * feat: the `error_type` classification vocabulary is threaded through the framework so `bifrost_error_requests_total` can carry a normalized fault-domain label alongside `status_code` (#7141)
  * feat: model pricing supports time-of-day peak and off-peak rates. `TableModelPricing` gains `off_peak_cost_multiplier` (a nullable float) and `peak_hours` (a JSON-serialized `PeakHoursSchedule` of recurring weekly windows using IANA timezone names, weekday numbers and half-open `HH:MM` intervals that may wrap past midnight), added by migration `add_time_of_day_pricing_columns` and aliased into the datasheet package so the JSON shape stays self-contained. The cost engine evaluates the schedule against the request start time and scales usage-based charges by the multiplier when the request falls outside every peak window, applied once in `computeCostFromInput` so every modality is covered; flat `CostPerRequest` and `SearchQueriesCost` fees are excluded, as is `AdditionalCost`, which is discounted independently through its own pricing rows. Both fields are exposed on `PricingPatch` in the OpenAPI and governance schemas, with `off_peak_cost_multiplier` bounded to `(0, 1]` (#6574, #6575, #6576)
  * feat: `SecretVar.RedactedIfSecret()` is used for non-credential fields in `ProviderConfig.Redacted()`, `Config.GetAllKeys()` and `Config.RedactMCPClientConfig()`, so regions, endpoints, service URLs and MCP connection strings surface as plaintext while anything env-var or vault-backed stays masked (#7085)
  * feat: GA realtime transcription sessions are routed, governed, logged and priced through the normal framework pipeline, resolving the routing model from the nested `audio.input.transcription.model` that arrives in `session.update` rather than from the connect URL (#7089)
  * feat: a `use_openai_endpoints` column on the provider keys table, added by migration `add_use_openai_endpoints_column`, opts a Bedrock key or alias into Bedrock's OpenAI-compatible endpoints instead of Converse (#7073)
  * feat: the reserved tool-namespace list a provider keeps for its own server tools is read from the datasheet row `reserved_tool_namespaces`, so a namespace collision check no longer requires a code change (#7084)
  * fix: `exchangeRefreshToken` includes `client_secret` only when the secret is non-empty, matching `exchangeCodeForTokensWithPKCE`. Public OAuth2 clients registered against servers that support only `token_endpoint_auth_method: none` have no secret, and unconditionally setting `client_secret=` sent an empty `client_secret_post` attempt that strict authorization servers answered with `invalid_client`, flipping the token row to `needs_reauth` even though the refresh token was valid (#7042)
  * fix: the virtual key `allowed_models: ["*"]` handling for governance added in #6767 is reverted, returning the wildcard-with-empty-synced-catalog case to its previous behaviour (#7053)
</Update>

<Update label="compat" description="0.3.0">
  * fix: removed the namespace-tool flattening that ran under `should_convert_params`; it spliced nested functions into the top-level list without a namespace prefix, so two namespaces sharing a function name produced duplicate tool names and an upstream 400. Flattening now lives in Bifrost core for every provider whose wire lacks the `namespace` type, with unique `<namespace>__<function>` names and response-side mapping back. `should_convert_params` is still accepted so existing configs load, but it no longer changes any request (#7048)
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="governance" description="1.8.0">
  * feat: MCP tool logs record governance entity names beside their IDs, so the plugin stamps `user_name`, `team_name`, `customer_name`, `business_unit_name` and the multi-valued team, customer and business-unit sets onto the log entry at ingestion instead of leaving the dashboard to render raw UUIDs (#7154)
  * feat: endpoint-attributed MCP inspections carry bounded identity (device, app key, server label, tool name, decision) from the gateway into the MCP authorization path, so an inspected tool call is attributed without trusting payload-supplied headers (#6959)
  * feat: governance errors are classified into the normalized `error_type` vocabulary, so a 429 raised by a governance rate limit is distinguishable from an upstream 429 on `bifrost_error_requests_total` (#7141)
  * feat: model allow and block lists accept `regex:` entries. The `*_patterns` fields added earlier in this release window were withdrawn in favour of the prefix form, which needs no new schema fields (#6988, #7133, #7134)
  * fix: `UsageTracker.Cleanup()` cancels and waits for the periodic reset worker before taking its final budget and rate-limit snapshots. It previously dumped first, so `trackerCancel()` could cancel an in-flight rate-limit dump and fail with `failed to dump rate limits to database: context canceled`, and the final dump was not guaranteed to be the tracker's last database writer. A queued ticker event can no longer start another reset cycle during shutdown, the current cycle stops when the tracker context is cancelled, and `context.Canceled` is treated as an expected result only when that context was actually cancelled (#7099) [@Constantine3](https://github.com/Constantine3)
</Update>

<Update label="jsonparser" description="1.6.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="logging" description="1.8.0">
  * feat: MCP tool log rows carry governance entity names alongside their IDs, written from the request context at ingestion, so nothing is resolved on read (#7154)
  * feat: endpoint-attributed MCP observations are written onto both the pending and the final MCP tool log entry, with the observation snapshotted so it cannot alias across async entries (#6959)
  * feat: cost recalculation honours time-of-day peak and off-peak pricing, scaling usage-based charges by `off_peak_cost_multiplier` when a request falls outside every declared peak window (#6575)
  * feat: GA realtime transcription sessions are logged and priced through the standard pipeline with transcription-aware pricing (#7089)
  * fix: a request whose caller disconnected before the upstream finished is logged and finalized deterministically rather than roughly half the time, so abandoned requests no longer leave a log row stuck in its pending state (#6972)
  * docs: clarified that `CountRecalcTargets` reads a materialized view that can lag, so its `Total` is an approximation rather than an exact count (#7078)
</Update>

<Update label="maxim" description="1.7.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="mocker" description="1.6.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="modelcatalogresolver" description="1.1.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="otel" description="1.5.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="prompts" description="1.1.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="routing" description="1.1.0">
  * fix: a continuation turn, such as a tool result following a prior user message, is classified instead of skipped when no active session state is found. `BuildInputWithDisposition` now returns the populated `ComplexityInput` including `LastUserText` for trailing-continuation turns in both the chat and responses paths, and `computeComplexity` falls back to classifying that recovered text, so new or recovered sessions get a tier assignment; the skip path applies only when `LastUserText` is also empty (#7122)
</Update>

<Update label="semanticcache" description="1.6.3">
  * chore: upgraded core to v1.9.0 and framework to v1.7.0
</Update>

<Update label="telemetry" description="1.7.0">
  * feat: `bifrost_error_requests_total` carries a normalized `error_type` label alongside `status_code`, drawn from a closed prefix-structured vocabulary (`caller_*`, `policy_*`, `provider_*`, `bifrost_*`, `_OTHER`), so alarm expressions can separate fault domains with a single regex instead of enumerating status codes (#7141)
</Update>
