Skip to main content
v2.2.6
Check config before upgrading. Fresh deployments now enforce auth on inference by default (enforce_auth_on_inference), and deployments with MCP OAuth discovery enabled must set oauth2_server_config.issuer_url or the gateway will not start.

Changelog

v2.2.6 moves the enterprise gateway onto OSS transports v2.2.6. Callers that use a plain virtual key are now attributed to the key’s holder, and their organization’s budgets and rate limits are enforced. The OSS release turns inference auth on by default for fresh deployments, requires issuer_url for MCP OAuth discovery, scopes the semantic cache per virtual key, and streams Claude tool-call arguments incrementally on Bedrock and Vertex.

🌎 Open Source Features

  • Inference Auth On by Default - enforce_auth_on_inference now defaults to true for fresh deployments when config.json leaves it out, file-only deployments included. Creating the first enabled admin also turns it on unless the request sets it explicitly. A stored database value always wins and an explicit false is always kept. To keep unauthenticated inference, set "enforce_auth_on_inference": false in config.json.
  • MCP OAuth Discovery Requires issuer_url - When mcp_server_auth_mode is oauth or both, oauth2_server_config.issuer_url must be set (env syntax env.MY_VAR works), or config load fails and PUT /api/config rejects the save. The issuer is no longer derived from the request Host header, and discovery responses carry Cache-Control: no-store. Set it before upgrading any deployment with MCP OAuth discovery enabled.
  • Provider Dial Target and Proxy Changes Need Real Auth - With dashboard auth off, changes to provider base_url, allow_private_network, key endpoint URLs, provider proxy, custom CA certs, skipped TLS verification, absolute request_path_overrides and the global proxy URL now return 403 unless the request carries a genuine admin credential.
  • Private Framework Config URLs Rejected - pricing_url, model_parameters_url and mcp_library_url are checked at save time and at dial time against private, link-local and CGNAT addresses, and redirects are not followed. Air-gapped setups should use file:// URLs.
  • Semantic Cache Scoped per Virtual Key - Cache buckets are partitioned by virtual key, so a shared cache_key or default_cache_key no longer serves one virtual key’s cached response to another. Entries written before the upgrade are not reused (expect a cold cache), and a per-request threshold override can only raise the configured threshold, capped at 1.0. Docs
  • zstd Decoder Window Cap - zstd-compressed request bodies whose frame header asks for a window above 100 MiB are rejected before any allocation.
  • Compat: Clamp Over-Limit Output Tokens - With should_convert_params on (UI: Convert Unsupported Param Values, or x-bf-compat: ["should_convert_params"]), a max_output_tokens, max_completion_tokens or max_tokens above the model’s catalog limit is lowered to that limit instead of being rejected by the provider. A thinking budget at or above the lowered cap is moved just below it, and each change is logged as a warning on the request.
  • Per-Message Effort Override - An effort-only system message ({"role":"system","content":[],"output_config":{"effort":"low"}}) now reaches Anthropic with the mid-conversation-output-config-2026-07-01 beta instead of being dropped. Models and providers without per-message effort (Vertex, Bedrock, Bedrock Mantle, Azure, DeepSeek, Fireworks, vLLM, SGL, OpenAI-shaped providers) drop it instead of failing, which fixes Claude Code output_config: Extra inputs are not permitted errors on Vertex.
  • Datasheet Control for Per-Message Effort - Per-message effort support can be set per provider and model with the datasheet field supports_mid_conversation_output_config, so a surface that ships the feature can be enabled without a release. Non-Anthropic providers also need beta_header_overrides: {"mid-conversation-output-config-": true} in their network config.

🐞 Fixed

  • Plain Virtual Key Caller Attribution - A request made with a plain virtual key is now attributed to the key’s holder, and the holder’s organization budgets and rate limits are enforced once, instead of only for access-profile callers. The holder’s team does not widen the tools the key can reach.
  • Claude Code Billing Header with Effort-Only Messages (OSS) - The Claude Code billing header is stripped again when an effort-only system message comes first, so guardrails no longer evaluate it and non-Anthropic fallbacks no longer receive it as prompt text. Anthropic attempts get the header restored in place.
  • Claude Tool-Call Argument Streaming on Bedrock and Vertex (OSS) - Tool arguments stream incrementally, so long Write calls no longer arrive in one burst and Claude Code no longer aborts with “Stream idle timeout”.
  • Tool-Result Cache Markers for gpt-5.6+ (OSS) - Anthropic cache_control markers on tool results become prompt_cache_breakpoint on Chat Completions and Responses for gpt-5.6+ on OpenAI, Azure, Bedrock and Bedrock Mantle, keeping the latest four.
  • Handler Panic Recovery (OSS) - A panic in a request handler returns a 500 and logs a stack trace instead of crashing the process.
  • Secret Redaction in Config Responses (OSS) - Env and vault-resolved values in provider alias configs and Bedrock endpoint overrides are masked in management API responses.
  • Admin Password Autofill (OSS) - Password managers no longer fill a saved host password into the new admin’s password field on the Security page.

📀 Base OSS version

transports/v2.2.6 (pinned as github.com/maximhq/bifrost/transports v1.6.12-0.20261006042654-8b4fce4f1709), with core v1.11.3, framework v1.8.1, governance v1.8.6, and logging v1.8.6.

🔌 If you are compiling plugin against this release - use following deps