v2.0.0
Changelog
v2.0.0 is the first stable release of the v2 line and aggregates everything shipped across the three v2.0.0 prereleases plus the final stabilization window. It introduces Bifrost Edge, a cross-platform device agent with fleet management, scoped approvals, and a kill switch run from the enterprise dashboard; a full guardrails stack (PII, secrets, and custom regex redaction, prompt-based classification, and MCP tool guardrails); an alerting system with declarative channels and CEL rules; a SCIM and OIDC identity sync overhaul that makes SCIM-owned identity authoritative; and cluster gossip v2 with dedicated typed streams for multi-node scalability. The base OSS release istransports/v2.0.0, which brings batch accounting, input/output cost split, overhead latency breakdowns, and a standalone routing plugin.⚠️ Breaking Changes
- Gemini tool preference flip - On the Gemini API surface, a request that carries both function declarations and Google Search without
include_server_side_tool_invocationsnow keeps the function declarations and drops Google Search. It previously did the opposite. Setinclude_server_side_tool_invocationsto send both (Gemini 3). Vertex is unaffected. HTTPTransportPreAuthHookfor Go plugin authors - Go plugins that implementHTTPTransportPluginmust add the newHTTPTransportPreAuthHookmethod. Credential injection (for examplex-bf-vk) must move there, becauseHTTPTransportPreHooknow runs after transport authentication. Compiled.soplugins that predate the method are skipped for that phase.- OTEL attribute rename - Bifrost-internal
gen_ai.*attribute constants are removed in favor of canonicalbifrost.*keys, and connectors emit the new names. Dashboards and alerts keyed on the old attribute names need updating. - Non-reversible database migrations -
merge_oauth_token_tables,drop_oauth_config_pkce_columns,drop_oauth_config_token_id_column, andadd_budget_reset_config_columnin the OSS base, plus two enterprise migrations, cannot be rolled back. See the Database Migrations section.
✨ Features
✨ Bifrost Edge
- Bifrost Edge - New Edge product line: enrolled devices route their AI traffic through Bifrost for policy enforcement, with device, MCP, and edge config management from the enterprise dashboard. Device-side details are covered in the Bifrost Edge changelog.
- Edge Fleet Management - Server-side device management: device inventory with per-device details, device login sessions that authenticate without a virtual key (with an optional virtual-key auth mode), app version tracking, and dedicated RBAC permissions for edge control. The devices page can be filtered by user, and device inventory sync is additive instead of replacing the stored set, so a partial sync no longer drops known devices.
- Scoped Approvals for Apps and MCP Servers - Edge Control approvals for AI apps and MCP servers can be scoped to specific teams or users instead of applying globally, with per-device overrides managed from the dashboard; enrolled Edge agents enforce the resolved scope on device.
- Scoped Kill Switch - The Edge kill switch can target a scope instead of the entire fleet; agents pick up the scoped state through inventory sync and enforce it locally.
- Interception Overrides Management - Scoped interception overrides are managed from a list in a dedicated sheet: search across existing overrides, add new ones on demand, and edit in place. A global approve or block decision states its effect and clears the scoped overrides it supersedes, and bulk removal asks for confirmation. Edge settings pages follow the standard config page layout.
- Server-Side Credential Issuance for Devices - Trust material for enrolled devices is issued and signed by the server. Devices no longer receive long-lived signing key material, cold-signing fails closed, the signing endpoint is rate limited, and each device carries a
remote_signing_capableflag so a fleet can be migrated in place. - Signed Agent Responses - Trust-relevant agent-facing responses are signed with an Ed25519 key that is independent of the interception key material, so an agent can detect a forged response even if the transport or a bearer credential is compromised.
- Encrypted Key Material at Rest - A migration re-encrypts any legacy plaintext private key found in the stored agent config, so key material saved by older releases is protected at rest.
- Separate Allowed Domains Configuration - Allowed domains are configured independently of the rest of the interception policy, so domain scope can be changed without touching other settings.
- Edge Agent Download Distribution - Edge agent builds are published to S3-backed download infrastructure through the release pipeline, with generated per-environment onboarding documentation and templates for rollout.
- License Management - Licenses are validated against a public key embedded in the binary at build time, with a license table, migration, and enforcement middleware. The Edge product has its own license validation, and
LicensePublicKeyis enforcement-only with keyless dev support.
🚨 Guardrails and Alerting
- Guardrails Redaction - New redaction pipeline for guardrails: detect, block, and redact actions for PII providers (Presidio and Azure Language PII, with multi-select entity search), secrets detection, and custom regex rules, with findings composed across guardrails into a single redaction result. Docs
- Redaction Modes and RBAC Reveal - Redaction supports logs-only and reversible modes, configurable per guardrail in the UI. Reveal of redacted log content is RBAC-gated, redaction and reveal are phase-scoped, guardrail replacements are published to trace exporters, and raw request/response payloads in extra fields are redacted when redaction is enabled.
- Streaming Output Redaction - Redaction applies to streaming output for PII providers, including Responses API streams.
- Prompt Guardrails - A new guardrail provider that classifies request and response content against a natural-language rule you write, with configurable model, output token ceiling, and timeout. It fails open on uncertainty by design, so only clear rule violations block. Prompt guardrail evaluations report their own token cost and debug output through
guardrail_debug. Docs - MCP Guardrails - Guardrail rules can target MCP tool traffic, not just model traffic, including redaction and transformation actions on MCP tool inputs and results, with backend config and a rules UI.
- Alerting - New alerting system with declarative channels and CEL-based rules: channel registry with delivery logic, an evaluation layer sourcing metrics from governance, alert history stored in the log store, config.json loading and reconciliation, a leader-lifecycle-driven alerting manager, a dedicated RBAC resource, and a full management UI with channel icons in history.
🆔 Identity and Access
- SCIM and OIDC Identity Sync Overhaul - Claim-driven role, team, and business unit sync is unified into a single funnel used by login, dashboard token refresh, and the periodic sweep. SCIM-owned users are frozen on OIDC login: an OIDC login no longer overwrites roles, memberships, or profiles that SCIM owns. A provider-wide
claims_sync_modesetting (“provisioning source”) replaces per-user provenance checks, manual memberships are adopted instead of deleted, admin-configured team-to-business-unit edges are protected from automated sync, and virtual keys are preserved when role profiles are re-applied. SCIM claim sync also remembers the last seen value per attribute, so a token that omits an attribute does not wipe state derived from it. Docs - SCIM Attribute to Access Profile Mappings - IdP attribute values can be mapped directly to access profiles, with schema support, validation and normalization, auto-assignment during import, role sync and recompute paths, and a mappings editor in the SCIM wizard. An existing override profile is preserved in place on re-login and role change, and a missing mapping attribute in a token is treated as no signal rather than an authoritative empty value.
- Wildcard and Glob Role Mappings - Attribute-to-role mappings in OIDC/SCIM configuration support wildcard and glob pattern matching on attribute values.
- SailPoint SCIM Provider - SCIM provisioning support enabled for the SailPoint identity provider.
- Okta
SyncAllUsersToggle - The Okta SCIM provider can sync non-active users as well, excluding suspended and deprovisioned ones, for organizations that stage users before activation. - Entra Provisioning Performance - Entra group and user fetches are parallelized and batched via the Graph API, with progress reporting, live import counters in the sync UI, and role filtering support.
- Keycloak Group and Role Propagation - Keycloak group names and roles are propagated to the idpUser during SCIM provisioning.
- Team/BU Mapping Ownership Management - OIDC team mapping ownership moves are transactional with preflight collision detection, the UI warns on team/BU mapping ownership moves and renames before saving SCIM config, business unit lookup uses
source_idwith name fallback and backfill, and OIDC-owned team and BU names are reconciled on mapping changes during login and sync. - SCIM Config Hot-Reload - SCIM provider configuration changes are gossiped cluster-wide so all nodes hot-reload without a restart.
- Service Accounts - Users get an
is_service_accountflag for non-human identities. Admins can create a service account without an email, service accounts cannot be used as a login identity, SCIM directory sync will not remove or edit them, and they are surfaced in the users table. - First-Time Admin Bootstrap Token - A one-time token flow creates the first admin user, replacing the previous manual bootstrap step.
- Audit Log Severity - Audit log entries carry a severity level, set through the audit middleware, so high-impact administrative actions can be filtered apart from routine ones.
👩💻 Platform and APIs
- Canonical
/api/governanceNamespace - RBAC, user, team, virtual key, access profile, business unit, SCIM and audit log routes now live under/api/governance. Legacy paths keep working through registered aliases, RBAC resource mapping follows the canonical paths, and the enterprise UI calls the new ones. Routing endpoints are extracted into a standalone/api/routing/*plugin with its own RBAC mappings. - Cluster Gossip v2 with Typed Streams - Nodes that advertise the
gossip:v2capability move governance, KV store, circuit breaker, load-balancing log, and cluster diagnostic traffic onto dedicated gRPC streams with independent lifecycles, instead of one shared stream. Broker mode gets the same separated lanes, anti-entropy runs on a 30-second interval with needs-full recovery when a sweep batch is missing, access-profile propagation broadcasts are coalesced, and per-send allocations are reduced. - Splunk Connector - New connector that exports logs and metrics to Splunk over HEC, with TLS client certificate support, an indexer acknowledgement pipeline for reliable delivery, request type in exported events, and a configuration UI.
- Inspect Endpoint - New
/inspectendpoint for pre-flight request evaluation. It runs the configured plugins without calling a provider, skips model and provider validation, and skips budget and rate-limit checks so an inspect call is never charged. - Token Exchange with SSO Application Credentials - MCP clients using
use_idp_credentialsreuse the SSO login application’s client id and secret, and those credentials are resolved unconditionally onto the token exchange IdP so Microsoft Entra ID style flows work without duplicate configuration. - Delegated MCP Token Exchange - Validated IdP tokens and OIDC sessions stamp an inbound bearer on the request context, and a SCIM-backed resolver wires delegated MCP token exchange to whichever SCIM provider is enabled.
- Cluster-Wide MCP Credential Cache Eviction - MCP OAuth token and per-user header credential cache evictions are broadcast cluster-wide, and credential grants are reconciled on user delete so a removed user loses access on every node.
- Cross-Instance MCP Connection State - A new node state store and heartbeat publish each instance’s per-client MCP connection state into the shared KV store, and an aggregate view compares them, so a client that is healthy on one node and unstable on another is visible instead of averaged away.
- MCP OAuth Refresh Worker - A cluster-gossiped refresh worker renews MCP OAuth tokens and triggers a reconnect hook, plus a
needs reauthgossip action that closes sessions requiring re-authorization. - MCP Tool Group Lookup by ID - MCP tool groups can be referenced by ID in addition to name during config reconciliation.
- Custom Branding - Logo and icon overrides are stored in a new enterprise branding table and served through
GET/PUT/DELETE /api/brandingplus an asset route, so the dashboard shell renders your brand instead of the default one. White-labelled deployments show a “Powered by Bifrost” attribution badge in the sidebar footer. - Quarterly Budget Reset for Enterprise Entities - Budgets on access profiles, teams, customers, and users support a quarterly reset duration in the UI, and budget reset configuration is editable after creation.
- Gateway Overhead Metrics in Connectors - Gateway-added latency (overhead) is exported to the BigQuery, Datadog, and Splunk connectors, and inference middlewares are wrapped with timing middleware.
- Connector Attribution Attributes - Connectors carry previously missing attribution attributes, with expanded coverage across the BigQuery, Kafka, Pub/Sub, and Datadog connectors.
- WebSocket Propagation Progress - Access profile propagation job progress is pushed over WebSocket events instead of polling.
- Optional Google Workspace Admin Email - Google Workspace
adminEmailis now optional; bulk sync is disabled when it is absent or the Directory API is unreachable. - Leader-Gated OAuth2 Sweep Worker - A leader-gated sweep worker purges expired authorize requests, revoked refresh tokens, and orphaned dynamic clients.
- Security Headers and
robots.txt- Enterprise bootstrap adds a security headers middleware and arobots.txtroute, and a skills orphan cleanup worker removes dangling skill records. - Access Profile Aware Virtual Key Resolution -
ensureUserVirtualKeyskips virtual key resolution when the user already has an access profile, removing an unnecessary lookup from the login path. - Enterprise Context Middleware - Every per-request fasthttp context is stamped with the enterprise marker through a dedicated middleware, so downstream plugins can rely on it being present.
- Responsive Enterprise UI - The dashboard adapts to smaller screens, page headers are consolidated into a single PageTitle component with a unified search and actions toolbar row, and filter-sidebar pages show a bordered main panel.
- User List Filters and Inline User Search - The users table can be filtered and sorted by role and filtered by identity type, and filter sidebars search users inline instead of loading the full user list.
- Enterprise Management Postman Collection - A generated Postman collection covers the enterprise management APIs.
🌎 Open Source Features
- Batch Accounting - Provider batch jobs are tracked in a new
batch_jobstable and settled asynchronously: per-model catalog batch rates on the results path, one idempotent aggregate cost log with the creating request’s identity, a background sweeper with ownership fencing, usage charged exactly once to the creating user’s budgets and rate limits, mixed-model repricing during recalculation, and a Batch Details block in the log detail view. - Claude-on-Vertex Batches - Vertex batch jobs route Anthropic models to
publishers/anthropic/...and build Claude-on-Vertex JSONL, round-tripcustom_id, and preservetools,toolConfig,cachedContent,labels, anddisplay_nameon Gemini and Vertex batches. - Input / Output Cost Split - Every log carries
input_cost,output_cost, andadditional_cost(guardrails, semantic cache, MCP) next to the total, across the relational store, ClickHouse, materialized views, recalculation, and the quota API; speech, transcription, and OCR carry cost as well, and the log detail view shows the split. - Bifrost Overhead Latency -
upstream_latencyandoverhead_latencyon every log, aggregated (avg, p90, p95, p99) in a new dashboard Bifrost Overhead chart. Overhead is decomposed by span self-time into serialization, conversion, plugins, middleware, key selection, queue wait, networking, client delivery, and scheduling buckets, persisted tooverhead_breakdown, and rendered as a stacked bar, with a Prometheus/OTEL histogrambifrost_overhead_latency_microseconds. - Routing Plugin - Routing rules and the complexity router are extracted into a dedicated
routingplugin that runs after governance so rules evaluate on the fully stamped context. Endpoints move to/api/routing/rulesand/api/routing/complexity-analyzer-configwith deprecated/api/governance/*aliases, and complexity routing reads the text of mixed text-plus-image turns. - Notification Center - Role-targeted dashboard notifications stored in the database, delivered over WebSocket, with a topbar tray and
GET/POST /api/notifications. - Topbar and Responsive Dashboard - A persistent topbar with page titles, theme toggle, links, user menu, and version, responsive layouts across all views, and version-skew detection with an auto-reloading upgrading screen.
- MCP Per-User OAuth - MCP clients can hold per-user OAuth credentials and per-user headers, configurable from
config.jsonand the UI, with a documented shared versus per-identity token lookup contract, anoauth_config.resourceparameter (RFC 8707), and virtual key and user filters on the OAuth grants and MCP auth session sidebars. - MCP Connection Lifecycle and Tool Discovery - Discovered MCP tools persist and resync uniformly across all client types through a hash-gated core callback, surviving restarts and propagating across a cluster. Reconnects are make-before-break, sticky-client static header updates pre-flight verify and swap onto the live connection, a failed enable parks at
Disabledfor retry, and the globaltool_sync_intervalhot-reloads. - Air-Gapped MCP Catalog -
mcp_library_sync_interval: 0disables catalog sync, andfile://URLs load the MCP server library from disk. - MCP Metrics - MCP metrics are exported through OTEL and the Prometheus telemetry plugin.
- MCP Log Redaction and Plugin Logs - MCP tool logs carry redaction mappings and plugin logs.
- Sarvam AI Provider - Sarvam AI added as a first-class provider with chat, text-to-speech, and speech-to-text support. Docs
- Wafer AI Provider - Wafer AI is supported as a provider.
- Runware Chat, Catalog, and Media Operations - Chat completions, streaming, and Responses via Runware’s OpenAI-compatible endpoint,
ListModelsfrom the curated catalog, image upscale via/v1/images/edits, image-to-3D and async 3D via/v1/videos, provider-reported per-task cost, and a raw/runware_passthroughroute. - Video Edits -
POST /v1/videos/editsfor prompt-driven edits, upscaling, and background removal on an existing video (bytes, URL, or provider video ID), on OpenAI and Runware. - JSON Image Edits -
POST /v1/images/editsaccepts JSON bodies (URL or base64 images, typed extra params) in addition to multipart. - ElevenLabs Sound Effects - Text-to-sound generation support via
/v1/sound-generation. Docs - OpenRouter Speech, Transcription, and Embeddings - Text-to-speech and speech-to-text through OpenRouter audio endpoints, and embedding models in
ListModels. - Grok on Bedrock Mantle -
xai.models route through theopenai/v1Mantle path. - Gemini 3 Thinking Levels - A per-model
thinkingLevelsupport table clamps requested levels to implemented rungs, andreasoning_effort: "none"sets the model’s floor level instead of zeroingthinkingBudget. - Gemini Server-Side Tool Calls - Gemini
toolCallandtoolResponseparts surface asweb_search_callitems with their own call IDs and queries, unmapped built-in tool types are preserved on the native round trip, and eachthoughtSignatureappears exactly once on replay. Docs - Datasheet-Backed Compatibility - Anthropic, Bedrock, Cohere, and Gemini request shaping (adaptive thinking, native effort, disable-reasoning, mid-conversation system turns, computer-use and text-editor tool generations, default max output tokens, tool validation) is resolved from model capabilities instead of hardcoded model-name checks.
- Reasoning Effort None - Models that reason by default but cannot reason with tool calls get
reasoning.effort: "none"instead of losingreasoningentirely. - Anthropic Default Fallback Routing - Anthropic’s
fallbacks: "default"preset is preserved through the Bifrost round trip, with the server-side fallback beta header injected for default-routing requests. - Mid-Conversation Tool Changes - The mid-conversation tool changes beta header is supported for Anthropic and Bedrock Mantle.
- Reasoning Token Tracking - Anthropic extended-thinking tokens are tracked as reasoning tokens across chat, responses, and passthrough.
- Adaptive Thinking on Raw Passthrough - For adaptive-only Anthropic models, a legacy
thinking.type: "enabled"block is rewritten to the adaptive form on the raw passthrough body as well as the typed request path. - URL Sources Inlined for AWS-Hosted Claude - URL-sourced images and documents are fetched and inlined on the native-Anthropic path, since Bedrock Mantle rejects URL sources. Fetches go through the SSRF-safe dialer with a size cap, and a failed fetch aborts the request rather than silently dropping an attachment.
- Typed Embeddings on Bedrock - Titan V2
embeddingTypesand Cohereembedding_typeson Converse, native invoke, and LangChainBedrockEmbeddings, with a typedEmbeddingData.EncodingFormatforint8,uint8,binary,ubinary, andbase64vectors. - Rerank Upgrades - Structured JSON documents,
return_documents,next_tokenpagination, caller document IDs preserved, Cohere-shaped errors, cross-provider responses converted back to the caller’s wire shape, and/genai/v1/rankserved cross-provider, with aninput_cost_per_querypricing field. - Bedrock Project Scoping - Optional
project_idin Bedrock and Bedrock Mantle key configs, with per-alias overrides for Bedrock, Bedrock Mantle, and Vertex, plus UI support. - Bedrock VPC Endpoints - AWS Bedrock keys can target VPC endpoints, keeping Bedrock traffic on private networking. Docs
- Bedrock HTTP/2 PING Keepalives - The Bedrock provider can send HTTP/2 PING frames on idle connections through
http2_ping_interval_in_seconds(0 disables), exposed in the provider network config UI, so quiet streams survive intermediaries that cut idle connections. Docs - Bedrock Batch Role ARN - A
batch_role_arnon Bedrock key config passes a service role to Bedrock batch jobs for S3 access, taking priority over anyrole_arnin the request. - OpenAI Ultrafast Service Tier -
service_tier: "ultrafast"is forwarded only to supporting models, billed at dedicated rates, with custom pricing override fields. - Service Tier on Logs - Logs record the tier actually served, including Anthropic’s
service_tierfrommessage_starton streams, with a Service Tier column and detail field; repricing uses the served tier. - Pricing Fields - Per-request flat fee (
cost_per_request) for models billed per call, megapixel image tiers, per-size and joint size-plus-quality image rates forgpt-image-1-style models, andinput_cost_per_queryfor rerank, flowing through datasheet sync, the cost engine, custom overrides, and the pricing override form. - Model Catalog Pricing - Pricing data added to the model catalog, and
/api/models/detailsexposes resolved pricing overrides with catalog rows resolving overrides server-side, so the catalog shows the price actually charged. - Virtual Key Budget Overrides - Temporary budget overrides add
override_amounton top ofmax_limitand run either for a fixed number of reset cycles or until removed, configured throughoverride_mode,override_cycles_total, andoverride_anchor_resetacross the database, governance store, admin APIs, and UI. - Per-Model Budgets and Rate Limits - Virtual key provider configs accept budgets and rate limits scoped to individual models, surfaced through a unified budget override manager that groups provider and model budgets together.
- Quarterly Budget Windows - Budgets support a quarterly reset period with a configurable fiscal start month for virtual key provider configs and the customer entity, and budget UI labels surface the configured fiscal year start.
- Budget Usage Reset Coverage - The reset budget usage flow covers teams, customers, model limits, and provider governance, not only virtual keys.
- User Scope for Routing and Pricing - Routing rules and pricing overrides can be scoped to individual users, with a
user_idCEL variable in routing rules and a user picker in the pricing overrides UI. - Async Webhooks - Webhook delivery for async jobs, with endpoints configurable through
config.json, the admin API, and the UI, an SSRF-safe dispatcher with retries, paginated delivery history, and inferencerequest_idpropagation through jobs and payloads. - Background Model Catalog Refresh - Each provider’s list-models response is re-fetched on a
live_models_sync_interval(default one hour,0disables), so models an upstream starts serving after boot appear without a restart. - Stream Truncation Detection - A new SSE truncation interface and EOF handling across providers surface upstream stream death as an error instead of a clean
[DONE]. - Trace Redaction - Phase-scoped redaction and revealing, a transient redaction data field for guardrails, and trace content redaction before connector export.
- Durable Background Jobs - New
sidekiqbackground-job table, store methods, and runner with recovery; cost recalculation migrated to a durable, resumable, cancellable job with polling instead of SSE, with partitioned job claiming and FIFO ordering per key. - Audit Log Object Storage - S3/GCS object storage config schema for audit log archival, with
archiveInterval,archiveGracePeriod, andarchiveMaxObjectBytessettings, plus a toggle to always retain request and response content regardless of retention cleanup. - Alerting Configuration Schema - Alerting schema in
config.schema.jsonwith declarative channels and CEL-based rules, plus Helm chart support. - Splunk Connector Configuration -
config.schema.json, Helm values, and dashboard entries for the Splunk HEC observability connector. - HTTP Transport Pre-Auth Hook - A new
HTTPTransportPreAuthHookplugin phase runs before transport authentication so plugins can inject credentials such asx-bf-vk, with avirtual-key-from-confignative plugin example. - Plugin Inject Limits - Per-plugin
semaphore_sizeandinject_timeoutonPluginConfigbound observabilityInjectcalls so a hung connector releases its slot. - Harness Session Autodetection - Claude Code, Codex CLI, and OpenCode session headers populate the session ID when
x-bf-session-idis absent. - Auth and Model Check Skip Paths - Context keys let trusted internal callers bypass auth resolution, and evaluate-only requests such as
/inspectbypass virtual key provider and model allowlists while budgets and rate limits still apply. - Passthrough Encoding Negotiation - Forwarded
Accept-Encodingis filtered to decodable codecs (gzip, deflate, brotli, zstd; gzip/identity for streams), and chained content encodings are decoded. - Dimension Scope Ceiling - Grouped log analytics (rankings, histograms, key pairs) are bounded to the customer, team, business unit, user, and virtual key ids the caller may see.
- Canonical Model Names - Dashboard model rankings show canonical model names instead of inference-profile IDs.
- OAuth2 Hardening - Allowlist for private-use redirect URI schemes (RFC 8252 section 7.1) and a
shouldSweepgate on the OAuth2 sweep worker. - Mirrored Schema Support -
schema_url/BIFROST_SCHEMA_URLfor mirrored schema locations in isolated deployments. - Vertex Single-Region Config - Single-region configuration is enforced in Vertex key config.
- Helm Chart Updates -
bifrost.alerting, audit-log object storage,postgresql.external.portstring support,bifrost.mcp.toolGroups[*].id, broker clustering viabifrost.cluster.type: brokerwith broker address, port, and TLS settings, external PostgreSQL for the logs store, and nodeSelector, tolerations, and affinity on hosted PostgreSQL. - Expanded OTEL Metric Attributes - Metrics carry a service instance id plus team, customer, and business unit ids and names, so exported series can be sliced per tenant without post-processing.
- Separate OTEL Metrics Pipeline - The OTEL collector supports a metrics tab independent of traces, with separate headers for traces and metrics.
- OTEL Export Timeout - A new
export_timeoutsetting (default 5 seconds) bounds how long a slow or unreachable collector can hold an export goroutine. - Throughput Metrics - Tokens per second histogram endpoints, dashboard metrics, and throughput in model rankings and trend data.
- W3C Trace ID Propagation - Requests carry a W3C trace id on the context, so gateway logs join cleanly with upstream traces.
- Grouped Logs View - The logs table groups fallback chains under expandable roots through a
roots_onlyfilter with child aggregates, and the model catalog persists tab, search, and provider in the URL. - User Agent and App Attribution in Logs - Logs and MCP tool logs record user agent, app, source, decision, app key, and device id, with custom user-agent mapping and dashboard dimension rankings; MCP tool logs observed by the Bifrost Edge agent can be ingested with attribution.
- Server-Side Tool Calls in Logs -
web_search_call,code_interpreter_call, and similar Responses items render their full payload in log detail. - Status Code Badges - Error and passthrough logs show the upstream HTTP status code in the log detail header.
- S3 Log Export Metadata - Additional metadata is written alongside S3 log exports.
- Matview Maintenance Off Switch -
matview_refresh_intervalaccepts"off"to disable log store materialized view maintenance entirely. - Database Connection Controls - New
conn_max_idle_time(default 5 minutes) on both config and logs stores, acache_ttl(default 60 seconds) for password-command credential resolution, and amatview_refresh_timeoutbounding a single refresh pass. - Routing Rule Validation - Routing CEL expressions and
scope_idreferences are validated at write time in create and update handlers. - Routing Info Headers - Routing info headers are emitted for streaming responses, inference and integration APIs, and error and passthrough paths.
- Access Profile Config Schema -
config.schema.jsonacceptsblacklisted_models(a denylist that wins overallowed_models), aweightseed for weighted routing, andmodel_budgetson access profile provider configs. - SSO Additional Scopes -
config.schema.jsonacceptsadditionalScopes, requesting extra OAuth scopes on top of the base set for authorization servers that gate claims such asgroups. - WebSocket Proxy Support - Realtime and Responses WebSocket connections route through the configured provider-level proxy (HTTP, SOCKS5, environment based) instead of always dialing direct.
- Configurable SCIM Buffer Sizes - A buffer size option on the HTTP client factory lets IdP token endpoints return headers larger than the 4KB default without failing SCIM and OAuth clients.
- Count Tokens Coverage - Count tokens support added for Bedrock Mantle, DeepSeek, and SGLang, plus a retrieve-stream method on the Responses API.
- Model Reasoning Metadata - A
ModelReasoningschema field and provider-qualified model id resolution for model parameter lookups, with a requiredmodelquery param and a 404 response ongetModelParameters. - Dashboard Export and Ranking Controls - A
RankingLimitfilter withallandlimitquery params, uncapped snapshots for PDF and CSV exports, per-tab export scope, and acache_hit_typesdashboard filter. - Async Entity Selectors - Teams, customers, and virtual keys load through async selector components instead of preloading full lists, and the customer list returns a server-computed virtual key count.
- User Assignment on Virtual Keys - Users can be assigned from the virtual key sheet.
- Connector Latency and User Email Export - Connectors receive Bifrost latency and overhead duration, and can export user emails.
- ChatGPT Passthrough - ChatGPT passthrough route on the OpenAI integration with dedicated request handling.
- Edge Control Fallback Pages - Fallback pages for Bifrost Edge control views (config, devices, inventory) backed by governance resolver support.
- Agent Handover Page - Agent handover page with seeded end-to-end data support.
- Shell Rewriter Hook - The UI handler exposes a
ShellRewriterhook for pre-hydration HTML rewriting.
🐞 Fixed
Enterprise
- /api/devices Auth Bypass - Stopped
/api/devicesbypassing auth via the/api/devprefix. - Governance for Inline Batch Requests - Budget and rate-limit checks run for every model in an inline batch request, not just the first.
- Model-less Request Budgets - Requests without a model now charge access profile budgets.
- Cancelled Request Accounting - Billed cancelled requests are counted in governance accounting.
- List Models Governance Checks - Budget and rate limit checks and usage tracking are skipped for list-models and other metadata calls, which do not consume model tokens.
- Access Profile Enforcement - Fixed model blocklist and key allowlist checks in access profiles, and allowed-provider narrowing for access-profile based flows.
- Multi-Batch Delta Sync - Every batch of a multi-batch governance delta send is applied, not just the first.
- Budget Lookups -
QuotaGovernanceForVKreads budgets and rate limits from the config store instead of a stale local store and propagates errors. - Per-Model Budget Cleanup - A user’s per-model budgets are deleted when the user is deleted.
- Virtual Key Auto-Attachment Removed -
ensureUserVirtualKeyno longer auto-attaches a virtual key on auth paths; users with an access profile skip virtual key resolution entirely. - Governance State Sync - Governance no longer blocks on state sync; requests are served from DB state while leader sync retries, and state-sync baselines are snapshotted under lock to prevent concurrent map read/write.
- Cluster Usage Sync CPU - Reduced CPU overhead of the cluster usage sync loop.
- Cluster Diagnostics Peer List - Built from the capability cache so it reflects live peers.
- Duplicate Job Execution - The sidekiq reaper was replaced with an atomic dispatcher, preventing duplicate job execution in multi-node clusters.
- Guardrail Streaming Headers - Headers are cloned and snapshotted so they are not dropped on streamed output.
- Guardrail Redaction Tool Results - Tool result text references are aligned before redaction, so redacted spans map back to the right content.
- Guardrails on Responses API - Instructions in Responses API payloads are extracted and transformed correctly.
- Prompt Guardrail Errors - Prompt guardrail failures return specific error messages instead of a generic intervention message.
- Redaction Shared References - Request/response objects are copied before redaction to avoid shared reference mutation; cloning only happens for logs-only mode.
- GraySwan Canonical Content - The GraySwan integration handles raw canonical chat content arrays, sends the correct trace id, and marks policy id as required for the Cygnal API.
- Datadog Plugin Environment Variables - Environment variable support added for fields that previously had to be set literally.
- Deprecated Connector Metrics - Removed deprecated metrics from connectors and updated Kafka and Pub/Sub metric names.
- Observability Limits - Limits are passed through to
SetObservabilityPlugins. - BigQuery Writer Double Close - Fixed a double close of the managed writer in the BigQuery connector.
- Keycloak Token Selection - The Keycloak auth cookie uses the access token so
realm_accessandresource_accessrole claims survive, and ID-token providers always use the encrypted ID token for session classification. - OIDC Session Token Split - Session storage separates the ID token from the access token with a backfill migration.
- OIDC Session TTL Floor - A minimum session TTL is applied in the token refresher.
- OIDC Token Endpoint - Explicit
tokenEndpointis preferred over the auto-constructed URL in OIDC config. - SCIM Token Rotation - Rotated session tokens are handled without dropping the session, and the SCIM inference middleware handles credential rotation for intercepted apps correctly.
- SCIM Group Listing - Non-SCIM memberships are excluded from the SCIM group list, so an IdP “push now” reconcile cannot silently adopt them as SCIM-owned.
- BU Mapping Reassignment - Business unit mapping ownership is transferred on SCIM group reassignment instead of erroring or duplicating.
- Identity Cache Key - The identity cache is keyed by email instead of token subject for stable resolution across tokens.
- Claim Enrichment - Claims are enriched from the provider on the token enrichment path, and department and title are filled from Keycloak claims when present.
- Token Refresh Race - Fixed a race condition in token refresh.
- User Attribution - Fixed user attribution on gateway request paths.
- MCP Caller Context - User name, email, team, and business unit are stamped uniformly for both virtual-key-authed and user-authed MCP callers.
- MCP OAuth Storage - OAuth flows and tokens migrated to
mcp_oauth_flowsandmcp_oauth_tokenswith auth-mode guards on DAC scopes, cascades, and reconciliation. - DAC Log Visibility - Row visibility is separated from org-identity disclosure; out-of-scope org fields on log rows are redacted instead of hiding the row.
- Notifications for Role-Authenticated Callers - Added the Notifications RBAC resource and forwarded NotificationStore methods through the enterprise config store wrapper.
- Large Payload Rejection - The request size threshold middleware runs before authentication so oversized payloads are rejected early.
- Migrations Before License Check -
LoadConfigruns migrations before the license check, so a fresh database no longer fails startup on a missing license table. - Base URL Normalization - Public base URLs are normalized consistently, and a base URL caching issue is fixed.
- Edge APIs on Postgres - Fixed Edge APIs when running on Postgres.
- Device and Background Job Stores - Fixes to the device config store and the durable background job store used by the device inspect flow.
- Device Inspect - Inspect no longer runs provider checks that could block it, and Responses instructions are handled correctly on the inspect path.
- Propagate Job Cancellation - Context cancellation is respected when acquiring the semaphore in the propagate job.
- Access Profile Broadcast - Removed redundant access profile change broadcast on update.
- Linked Scopes Cleanup - Deleting a linked scope now deletes the linked rule.
- Prompt Logging - The actual prompt is no longer logged back in responses.
- Pangea Removal - Removed the Pangea integration.
- Security Hardening - Fixed code scanning and threat-vector findings across the SCIM discovery proxy, virtual key resolver, proxy paths, and device signing endpoints, including leaf-sign rate limiting and signature checks.
- UI Fixes - Long text in user-group columns truncates with tooltips, the duplicate “Apply on” section in the CEL rule sheet is removed, the user detail sheet uses the standard virtual key selector, Edge Control query cache invalidation works, the sync users sheet can be closed during the importing step, SCIM wizard save-time validation errors are routed to the step that owns them, and the overrides sheet matches the device details sheet width.
Open Source
- Structured Output Schema Order -
response_formatJSON schemas are forwarded byte-for-byte to OpenAI, Anthropic, Bedrock, Gemini, and Cohere so fields generate in the caller’s declared order. - Path Normalization Auth Bypass - Fixed a path normalization flaw that allowed auth to be bypassed.
- Connector Header Redaction -
Authorization,x-api-key, Cloudflare Access, and AWS ALB OIDC headers are redacted before export to every observability backend. - DAC-Scoped VK Reads -
from_memoryvirtual key reads are blocked for DAC-scoped callers. - Anthropic Compaction Token Undercounting - When Anthropic returns
usage.iterationsfor a compaction pass, compaction iteration tokens are folded into billing paths instead of only the reply pass being counted, fixing a large output token undercount. - Anthropic Server-Side Fallback Tokens - Fixed fallback token computation for Anthropic server-side calls.
- HTTP 529 Rotating Credentials - Anthropic
overloaded_erroris treated as a transient server error; the same key is retried with backoff instead of being rotated away. - Anthropic Fallbacks and Billing - Fallback handling and refusal responses on the Anthropic surface are fixed, and billing attributes usage to the fallback model actually served.
- Anthropic Tool ID Sanitization -
tool_use/tool_resultids are sanitized to Anthropic’s charset. - Anthropic Mid-Conversation System Messages - A system turn that cannot be forwarded natively is inlined as a user turn instead of being dropped.
- Anthropic tool_search - Server-side
tool_searchis forwarded and rebuilt on the Responses path, tool search types are normalized, and server-side tool invocation opt-in reaches the Gemini declaration-drop gate. - Anthropic Costing - Corrected inference geo cost and cache rate for fast mode.
- Encrypted Reasoning Handling - Replayed encrypted reasoning no longer mints a mismatched item id, an upstream 400 on unverifiable content strips the reasoning and retries once (covering
/v1/responses/compactand count-tokens requests, and recognizing Anthropic’sredacted_thinkingrejection), and Cohere emits encrypted reasoning alongside the summary rather than instead of it. - Reasoning Replay on Chat-Shaped Requests - Fail-soft strip of replayed reasoning on
reasoning_detailshandles chat-shaped requests, not only Responses-shaped items, so mid-conversation model switches no longer surface “Invalidsignatureinthinkingblock”. - Thinking Signatures on Responses Content Blocks - Signatures are stripped off content blocks, not just
encrypted_content, and only reasoning items are dropped when nothing survives. - Reasoning Content Rejected by OpenAI and Azure Models -
reasoning.contentis no longer sent to non-gpt-oss reasoning models;summaryandencrypted_contentcarry everything those models accept. - Thinking Block Typing on Streams - Reasoning items with both an encrypted payload and a visible summary open as
thinkingblocks instead ofredacted_thinking. - Redacted Thinking Round-Trip - Anthropic
redacted_thinkingblocks round-trip on the Responses surface. - Replayed Thinking Blocks via
bedrock/Prefix - Content-lesstool_resultblocks are kept, interleaved block order is preserved,incompletemaps toerroron Converse, and pending reasoning is consumed by its owning item, so multi-turn tool use no longer wedges. - Grok Reasoning Effort - A substring match on “grok-3-mini” made newer Grok models silently lose
reasoning_effort; it is replaced with an exact-match deny-list that normalizes routing prefixes and suffixes. The shared OpenAI-dialect normalizer also no longer downgradesxhightohighbefore the xAI compat pass. - Minimal Reasoning Effort on GPT-5 Models -
reasoning_effort: "minimal"is preserved for GPT-5-family models instead of being downgraded tolow. - DeepSeek Thinking on Multi-Turn - Thinking is no longer silently disabled for ordinary multi-turn conversations through the OpenAI-compatible surface.
- Empty Structured-Output Streams - Content events are emitted when a tool-based structured-output call is reassembled on the Responses streaming path, affecting Vertex, Bedrock Mantle, and Azure Claude.
- Bedrock Reasoning and Cache Control - Double emission of reasoning content on Bedrock streams is fixed,
cache_controlmarkers translate through invoke and Converse paths, tool ordering intoolConfigis deterministic for prompt cache hits, and reasoning blocks with an absent text key are no longer sent. - Bedrock Streaming Correctness -
ConverseStreamreportsstopReason: tool_usefor tool-use turns, andmessage_startcarries an all-zero usage object when figures are unknown so strict clients accept the frame. - Bedrock Content Retention - InvokeModel decodes Anthropic type-discriminated image, tool use, and tool result blocks instead of dropping them, document-only messages are accepted, and office and PDF documents sent as OpenAI
type: "file"work. - Bedrock Header Signing Isolation - Caller headers stored for Anthropic OAuth passthrough are no longer forwarded to other providers, preventing SigV4 signature mismatches.
- Bedrock Tool Use IDs - IDs over 64 characters or outside Bedrock’s charset (for example Gemini thought-signature IDs) are aliased deterministically on both
tool_useandtool_result. - Bedrock Stop Reasons -
content_filterandguardrail_intervenedstop reasons map toincompletestatus with acontent_filterreason. - Bedrock Stop Sequences for Nova and Titan - Bedrock Converse camelCase
stopSequencesmaps to the neutralstopparameter; 81 catalog rows were silently losingstop. - Bedrock Truncation Signal -
max_output_tokenstruncation is signaled on the Responses API. - Bedrock Reasoning Config -
reasoning_configis preserved on cross-provider translation so fallbacks keep extended thinking. - Bedrock Error Type - The AWS exception type (
X-Amzn-Errortype) is surfaced on non-streaming Bedrock error responses instead of being dropped. - Bedrock Mantle Streaming - Registered in
ProviderSendsDoneMarkerso streams end afterfinish_reason, andservice_tieris dropped for Bedrock Mantle instead of forwarding a field it rejects. - Gemini 400s on Claude Code Traffic - Trailing assistant prefills are trimmed and mid-conversation system turns are inlined for Gemini and Vertex, and
extra_fieldsare echoed on/anthropic/v1/messages. - Gemini Tool Preference - When tool combination is disabled, function declarations win over Google Search so the model can still call the caller’s tools (see Breaking Changes).
- Vertex Mixed Tools - Vertex AI accepts function declarations and Google Search in the same request without
includeServerSideToolInvocations, andretrievalConfig.latLngis preserved. - Gemini and Vertex Fidelity -
generateContentkeepscandidates[0].safetyRatingsandavgLogprobs, truncated responses reportMAX_TOKENSinstead ofOTHER, valid integer constraints in tool schemas are accepted, and Vertex cached-content methods honour API key or context header auth. - Gemini Grounded Streaming - The web-search flag is reset when recycling pooled stream state so
web_search_callitems keep emitting. - Gemini Fixes - Web search options map to Google Search grounding, file upload MIME types are preserved, and video reference fields map to instances.
- URL-Sourced Files and Images -
gs://URIs are forwarded to Gemini and Gemma asfileData.fileUriand read from Cloud Storage for Claude-on-Vertex,s3://references go to Bedrock Converse as ans3Locationsource (skipping the 25 MiB inline cap), Bedrock rerank synthesizes the foundation-model ARN from a bare model ID, OpenAI file blocks keepfile_url, and non-http schemes pass through on OpenAI and native-Anthropic paths. - GenAI SSE Heartbeats - GenAI streams delimit heartbeat comments so Google SDK clients preserve the following event, while older openai-go clients keep the bare heartbeat.
- SSE Heartbeat Corruption and Compatibility - The stream reader will not emit a heartbeat mid-line, and the heartbeat frame no longer carries a trailing blank line that made some SSE decoders abort mid-stream.
- Proactive SSE Disconnect Detection - Client disconnects during streaming are detected proactively instead of only when a producer loop attempts a write, fixing false-success logging on fast upstreams.
- Closed Channel Panic on Stream Shutdown - Fixed a race where a heartbeat goroutine mid-send at shutdown could panic with “send on closed channel”.
- Empty Stream Nil Channel - Stream requests return a closed non-nil channel for empty streams instead of
(nil, nil), which previously hung consumers on a nil-channel receive. - Stream Termination Edge Cases - A nil delta paired with a non-nil finish reason no longer aborts the stream, and GPT-5-series detection tolerates prefixed model names.
- Null Tool-Call Function Name on Streaming - Streaming continuation deltas no longer materialize an absent tool-call function name as
null. - Streaming Accumulation - Citation annotations and
finish_reasonare preserved in the accumulated streaming response. - Streaming Error Panic - Nil-safe tracing span lookup prevents panics on streaming errors.
- Azure Responses Stream Errors - Terminal
errorandresponse.failedevents inside an open HTTP 200 SSE stream are surfaced as errors with nested type, code, and message. - Azure Auth Headers - Azure auth headers are passed in helpers.
- HuggingFace Streaming Usage -
stream_options.include_usagedefaults on chat streaming, so streamed calls stop reporting zero tokens and zero cost. - HuggingFace Model IDs - Backfilled HuggingFace model ids no longer duplicate the inference-provider segment.
- vLLM Responses Streaming - vLLM responses-stream chunks and completion events are forwarded instead of silently discarded, and truncation is handled.
- OpenCode max_tokens -
max_tokensis preserved for OpenCode-compatible chat endpoints, and OpenCode Responses requests forward directly to/v1/responses. - OpenAI Responses Input -
roleis stripped from non-message input items and compaction requestinputis serialized correctly. - additional_tools Support -
additional_toolsmessage type support added, preserving nested tool types on/v1/responses. - OpenAI Parameters - Service tier honored in chat completion and max reasoning effort capped.
- Realtime Transcription Sessions - GA transcription-type sessions supported in
POST /v1/realtime/client_secrets, andresponse.createinput is guarded. - Diarized Transcription -
diarized_jsonsegments and ElevenLabs speaker passthrough supported. - Transcription Filename Dropped - The client’s multipart filename is carried through transcription ingress, so non-WAV containers are no longer relabelled and rejected upstream.
- WebSocket Writes After Disconnect - A broadcast racing a client disconnect could panic on a nil connection or deliver to an unrelated client’s socket; clients carry an explicit closed flag and a close that blocks until in-flight writes finish.
- Realtime Heartbeat Panic on Disconnect -
stopHeartbeatwaits for the heartbeat goroutine to exit; a ping on a recycled connection previously had no recover and took down the whole process. - MCP Reconnect and Lock Ordering - A lock-order inversion in the connection checker is broken, ephemeral clients are rebuilt across the whole connect and init retry, last-known tool maps survive close-first reconnects, and background reconnects are deduped.
- MCP OAuth Session Correctness - Reauthorize is restricted to shared OAuth clients, inactive tokens are rejected on validation, the OAuth flow claim is atomic against concurrent reauth, stored scopes survive decode failures, and a verify-headers double-submit race is closed.
- MCP Tool Errors Replayed as Success - Failed MCP tool executions are marked as errors instead of being replayed to the model as successful results.
- MCP Tool Sync Interval Corruption - The enable/disable toggle no longer corrupts
tool_sync_interval, negatives are rejected instead of silently disabling sync, and re-enabling a per-call client restarts its discovery cycle. - MCP Tool Map Staleness -
SetClientToolsreplaces the in-memory tool map instead of merging, so tools removed upstream leave memory. - MCP SSE Reconnect Identity -
OnConnectionLoston SSE MCP clients is gated on connection identity so a stale connection cannot tear down its replacement. - MCP Tool Ordering - Deterministic MCP tool ordering for prompt cache stability.
- MCP Timeout Placeholder - The MCP tool execution timeout placeholder shows the real global default.
- MCP Inline-Auth Links - Callers are warned not to truncate the
#t=temp-token fragment. - Session Stickiness Reconciliation -
needs_session_stickinessis pinned acrossconfig.jsonreconciliation, so an unrelated file edit cannot revert a client to per-call. - Credential Cache Cancellation - Credential and user token cache fills propagate context, so a cancelled request unblocks instead of waiting on an unrelated leader, and versioned LRU entries prevent a stale read from evicting a fresh value.
- Budget Counters Reset on Force-Sync -
config.jsonforce-sync no longer overwrites live usage, last reset, and rate limit counters with file values. - Calendar Alignment Semantics - Enabling calendar alignment preserves the currently open window and applies from the next period instead of truncating in flight.
- Governance Rate-Limit Reset CPU - Guards against invalid reset timeouts, parallelizes resting-budget flows only when required, fixes the calendar-based alignment qualifier, and corrects override counts for multinode setups.
- Governance List-Models Call - Budgets and rate limits no longer trigger a list-models call.
- Budget Pruning Crash - Pruning tolerates missing records for cascade-deleted budgets and configs, fixing a startup crash for API-created model configs absent from
config.json. - Virtual Key Provider Bulk Replace - Provider config replacement is a single bulk operation instead of per-provider round trips, removing a hot-path slowdown at scale.
- Wildcard allowed_models Repair - Bare wildcard
allowed_modelsrows that broke admin provider updates are repaired. - Masked Key Persistence - Masked provider key previews are never persisted to config storage.
- Provider Key Name on Update - A key PUT that omits
nameno longer clears it, and already-exists errors keep constraint detail. - API Key Provider Selection - Fixed provider selection for API keys, and key selection is skipped on the anthropic provider with stale URL-path and direct-key context cleared.
- Passthrough Virtual Key Attribution - Passthrough calls via the Azure
api-keyheader now attribute to the virtual key. - Rerank for Custom Providers -
/v1/reranknow works with custom OpenAI-compatible providers. - Together and Alias Pricing - The management catalog resolves the runtime
togetherprovider to the datasheet identity, configured aliases price through their target model, and USD cost ticks for xAI usage are fixed. - Responses Stream Usage - Stream usage is persisted when providers omit or reuse sequence numbers.
- Log Count Accuracy and Matview Scope - The hybrid matview count no longer over-counts boundary buckets, and customer and business unit columns are added to the matview scope projection so team-data scope resolves without column errors.
- Lost Log Rows on Shared Trace IDs - Concurrent requests inheriting the same W3C trace id no longer overwrite each other’s pending log entry.
- Hybrid Log Token Usage - Token usage is rebuilt from denormalized columns in hybrid log list.
- Live Reload Model List - Provider reload no longer wipes the live model catalog before refetching, so a transient list-models failure cannot empty it.
- Model Discovery - Disabled keys are skipped when scheduling model-discovery fetches.
- Log Store Migrations - Removed a duplicate materialized-view rebuild step from the log store migration registry and fixed the app-column step running the wrong migration function.
- Config Store Migrations - Cleaned up the sidekiq table creation migration.
- Pooled Object Hygiene - Pooled ChannelMessage references are zeroed on release and orphaned deferred spans are swept in trace store TTL cleanup.
- Redis Vector Store TAG Escaping - All RediSearch special characters are escaped in TAG query values.
- Plugin Stream Errors - Structured plugin stream errors are emitted on integration routes.
- Telemetry - Request id and trace id forwarded, metrics cardinality explosion risk reduced, and status codes sent on OTEL metrics.
- SecretVar Parsing -
SecretVarJSON withref/env_varfields parses even whenvalueis absent. - Entra OBO Scope -
offline_accessis combined with the audience default scope for Entra on-behalf-of instead of replacing it. - OpenShift Arbitrary UIDs - Build-time group-0 ownership with no runtime chown.
- HTTP Server Timeouts - Bounded server timeouts and a request body limit are configured.
- Stream Delta Schema -
ExtraContentadded toChatStreamResponseChoiceDelta. - Dashboard - Active time period preserved when applying dimension filters, bucket size thresholds adjusted for month-range durations, user popover with
preferred_usernamefallback, provider-level keys filtered from the prompt manager selector, password validation skipped for redacted credentials, andModelMultiselectempty and error states. - Dashboard Sidebar - Removed unused sidebar icon imports that broke the UI build.
- pprof Content-Type - pprof endpoints set
application/octet-streamfor scraper compatibility.
🗄️ Database Migrations
Enterprise (config store), across the v2 line:- New tables for the Edge product (device management, agent auth, agent settings, scoped approvals, agent policy attempts), licensing (
enterprise_license), branding, alerting channels and rules, durable background jobs (sidekiq), OAuth2, and cluster node heartbeats, plus RBAC resources for edge control, kill switch, alerting, and notifications. - Column additions: guardrail rule
target, audit logseverity, deviceremote_signing_capable, useris_service_account, and access profile enhancements. - Forward only (cannot be rolled back):
ent_migrate_legacy_plaintext_agent_ca_key(re-encrypts a legacy plaintext private key; the plaintext value is deliberately not restored on rollback) andent_split_oidc_session_auth_token_column(splits the stored OIDC session token into separate id token and access token columns).
transports/v2.0.0 release notes; framework v1.6.0 alone carries 13 (7 config store, 6 log store). Two points matter for planning the upgrade:- The log store migrations alter
logsandmcp_tool_logs, the two highest-insert tables, and build several indexes on them. Run the upgrade during a low-activity window or expect elevated log-write latency while they run. merge_oauth_token_tables,drop_oauth_config_pkce_columns,drop_oauth_config_token_id_column, andadd_budget_reset_config_columncannot be rolled back. Take a database backup before upgrading.
🐙 Closed OSS Issues
- #123 - Files API support
- #2347 - MCP tool ordering is non-deterministic, breaking prefix-based prompt caching
- #3455 - Segfault/nil dereference panic in Bedrock provider
- #4215 - HuggingFace models show provider ID twice in
/v1/models, which breaks requests - #4318 - allowed_models persisted as bare ”*” string blocks subsequent provider updates
- #4353 - config.db corruption from masked-key preview in provider_configs JSON column
- #4367 - Image incompatible with OpenShift arbitrary UIDs
- #4402 - Vertex provider drops image blocks whose URL uses gs:// scheme
- #4477 - Passthrough calls using a Virtual Key log as actual key
- #4679 - Bedrock Responses API does not signal max_output_tokens truncation
- #4689 - Custom providers cannot set budget
- #4712 - ElevenLabs sound effects (/v1/sound-generation)
- #4780 - Anthropic server-side tool_search results are dropped on /v1/responses
- #4834 - /v1/rerank is not available with custom providers
- #4846 - Responses stream usage present in response.completed but not persisted in LLM Logs
- #4851 - Governance rate-limit reset causes high CPU in BumpRateLimitUsage
- #4870 - Pooled ChannelMessage retains request body, context, and undelivered response while idle
- #4940 - Show canonical model names instead of Bedrock inference-profile IDs in Model Rankings
- #4963 - Streaming finish_reason dropped from the accumulated (logged) response
- #5002 - gpt-4o-transcribe-diarize transcription fails due to string segment IDs
- #5010 - Server-side SSE keepalive to keep long-idle streams alive through intermediaries
- #5013 - OpenAI /responses/compact input serialized as a JSON object causing 400
- #5026 - Toggling an MCP client’s enable/disable switch corrupts its tool_sync_interval
- #5027 - MCP Tool Execution Timeout placeholder shows 0 instead of real global default
- #5036 - Plugin StreamInterceptionError is flattened on integration routes
- #5037 - Disabled keys break provider model discovery
- #5051 - Add Sarvam AI provider (chat + TTS/STT)
- #5061 - Streaming responses drop citation annotations from the accumulated message
- #5074 - Fallback routing model selection is truncating model names
- #5093 - Streaming /v1/responses drops Anthropic redacted_thinking blocks
- #5097 - Anthropic rejects replayed tool_use/tool_result ids from non-conforming upstream providers
- #5100 - additional_tools loses nested tool types on /v1/responses
- #5101 - Chat-to-Responses tool replay sends role on function_call input items
- #5108 - Bedrock reasoning_config silently dropped on cross-provider translation
- #5113 - Gemini/Vertex streaming stops emitting web_search_call items after first grounded request
- #5186 - Anthropic-surface replay of OpenAI encrypted reasoning mints a fresh item id and OpenAI returns 400
- #5206 - Bedrock ConverseStream reports stopReason=end_turn for tool-use turns
- #5211 - Bedrock streaming can drop with “unexpected EOF” when an intermediary severs a quiet stream
- #5256 - Concurrent HTTP requests sharing a W3C trace ID lose LLM log rows
- #5279 - OpenAI /v1/responses to Anthropic drops the tool_search_tool_regex type
- #5308 - Responses API image blocks missing required “detail” field when converted from non-OpenAI providers
- #5329 -
/api/logsreturns an incorrecttotal_countfor time ranges of 24 hours or longer - #5432 - Add TTS and STT support for OpenRouter
- #5433 -
/genaiendpoint rejects validminLength/maxLengthin tool schemas - #5472 - Bedrock rejects office and PDF document uploads via OpenAI
type:"file" - #5504 - vLLM streaming Responses API hangs forever and chunks are silently discarded
- #5546 - Upstream SSE stream death swallowed into a clean
[DONE] - #5551 -
transports/bifrost-http/libtest package does not compile on dev - #5552 - Refresh the live model catalog in the background
- #5554 - Provider reload wipes the live model catalog before refetching
- #5555 -
*StreamRequestreturns(nil, nil)for empty streams, so consumers hang forever - #5670 - Transcription drops the client’s multipart filename
- #5679 - Anthropic Messages does not propagate Gemini mixed server and client tool opt-in
- #5843 - generateContent drops
candidates[0].safetyRatingsandavgLogprobson Vertex AI responses - #5871 - AWS Bedrock Mantle streaming is broken
- #5874 - SSE heartbeat frame aborts streams for openai-go ssestream consumers
- #5885 - v1.6.8 omits message_start.message.usage on Bedrock-backed providers
- #5887 - DeepSeek thinking silently lost on all multi-turn requests via OpenAI-compat inbound
- #5890 - Chat completions surface drops tool_result
is_error - #5900 - Streaming continuation chunks materialize omitted tool-call metadata as null
- #5902 - service_tier silently dropped for gpt-5.4 family
- #5905 - v1.6.8 raw passthrough heartbeat can split SSE data lines and corrupt JSON
- #5925 - config.json force-sync overwrites budget current_usage and last_reset on startup
- #5978 - Gemini reports truncated responses as FinishReason OTHER
- #6044 - normalizeOpenAIReasoningEffort maps ‘minimal’ to ‘low’ for all OpenAI models
- #6240 - GenAI SSE heartbeat framing causes @google/genai to silently drop the following data event
- #6248 - OpenRouter embedding models missing from Semantic Cache dropdown
- #6334 - Gemini/Vertex provider fails on Claude Code assistant prefills and mid-conversation system turns
- #6342 - Anthropic ingress with bedrock/ prefix restructures replayed thinking blocks, wedging multi-turn tool use
- #6416 - Provider key update silently clears “name” when omitted, then the unique-name index 409s subsequent updates
- #6457 - OpenCode chat endpoints drop max completion limit
📀 Base OSS version
transports/v2.0.0 (pinned as github.com/maximhq/bifrost/transports v1.6.12-0.20260826193051-e4a30d6041c0)🔌 If you are compiling plugin against this release - use following deps
The enterprise repo is a multi-module workspace; thegithub.com/maximhq/bifrost-enterprise/* modules at v0.0.0 resolve via the replace directives to the release source checkout.
