Skip to main content

Overview

Bifrost exposes Prometheus metrics via two methods:
  1. Pull-based (Scraping): Traditional /metrics endpoint that Prometheus can scrape
  2. Push-based (Push Gateway): Push metrics to a Prometheus Push Gateway for cluster deployments
For multi-node deployments: Use the Push Gateway method to ensure accurate metric aggregation. Traditional scraping may miss nodes behind load balancers.

Pull-based Scraping

Bifrost automatically exposes a /metrics endpoint when the telemetry plugin is enabled (enabled by default). No additional configuration is needed.
When Bifrost’s authentication is enabled (auth_config.is_enabled = true), the /metrics endpoint requires credentials. You can authenticate the scraper with either the admin Basic auth credentials (admin_username / admin_password from your auth_config) or, on Enterprise, a Bifrost API key with the Metrics permission. Without valid credentials, Prometheus receives 401 Unauthorized responses and scraping silently fails.

Prometheus Configuration

Add Bifrost to your Prometheus prometheus.yml:
If Bifrost authentication is enabled, add basic_auth to your scrape config:
Prometheus scrapes over plain http by default, which sends the Basic auth credentials or API key in cleartext. When Bifrost is served over TLS, set scheme: https (and any required tls_config) in the scrape config so credentials are not exposed in transit.

Authenticating with an API Key Enterprise

On Enterprise deployments, you can scrape /metrics with a Bifrost API key instead of the admin Basic auth credentials. Create an API key with the Metrics permission (included in all default roles) from Settings → API Keys (see Creating API Keys), then pass it as a bearer token in your scrape config:
Older Prometheus versions that lack the authorization block can use bearer_token instead:
API-key auth for /metrics is an Enterprise feature. On the open-source build, the /metrics endpoint accepts only Basic auth (the admin_username / admin_password above).

Endpoint

Returns metrics in Prometheus exposition format.

Push-based (Push Gateway)

For multi-node cluster deployments, the Prometheus plugin pushes metrics to a Prometheus Push Gateway. This ensures all nodes’ metrics are captured regardless of load balancer routing.

Configuration

Basic Auth Configuration


Setup

  1. Navigate to Observability → Prometheus in the Bifrost UI
  2. The /metrics endpoint is shown at the top for scraping configuration
  3. To enable Push Gateway:
    • Enter the Push Gateway URL
    • Configure Job Name and Push Interval as needed
    • Optionally set a custom Instance ID
    • Enable Basic Authentication if required
    • Toggle Enable Push Gateway on
    • Click Save Prometheus Configuration

Available Metrics

The following metrics are available from both the /metrics endpoint and Push Gateway:

HTTP Metrics

The path label contains the matched route template (e.g. /genai/v1beta/models/{model:*}, /v1/messages/batches/{batch_id}), not the raw URL path. This keeps metric cardinality bounded by the number of registered routes instead of growing with every model name or resource ID that appears in a URL. For per-model breakdowns, use the model and provider labels on the bifrost_* metrics.

Bifrost LLM Metrics

Error Types

bifrost_error_requests_total carries two error dimensions. status_code is the raw fact. error_type is the normalized interpretation: a closed, low-cardinality vocabulary that answers whose fault was it, which is what alarms actually need. Values are prefixed by fault domain, so a success-rate alarm that should ignore caller mistakes is one clause rather than a list of reasons that grows over time:

Vocabulary

The set is closed, so this list is exhaustive. Declared values are stated by the code that produced the failure — it knew the reason, so there is no guessing. Inferred values are derived from the upstream’s status code, because a provider’s own error vocabulary cannot be trusted to mean the same thing twice. Three values are both. caller_cancelled and provider_timeout are named by the code that recognises them and are also implied by status 499 and 504. bifrost_internal is named on the paths that recognise themselves and is otherwise the fallback for a Bifrost-origin failure that named nothing.
Alarm on _OTHER directly. It means no producer declared a reason and no rule inferred one. The name is the OTel catch-all, matching the error_type label on the MCP metrics. It should be flat at zero; a rise means failures are going unclassified — most likely a new refusal reason was added somewhere without declaring its type — not that a new kind of failure is benign.

What the distinction buys you

The same status code means different things depending on who produced it, and error_type is what separates them:
  • A 429 is policy_rate_limited when a governance limit refused the request and provider_rate_limited when the upstream did. The first means your own limits are too tight; the second means you need more upstream capacity.
  • A 403 is policy_model_blocked when a grant refuses the model and provider_auth_failed when the upstream rejects the key.
  • A 503 is bifrost_dropped when Bifrost shed the request under queue pressure and provider_overloaded when the upstream is at capacity.
Classification always prefers what Bifrost knows over what a status code suggests, so a decision Bifrost made itself is never re-attributed to the provider. The same value is stamped on the provider-attempt span as bifrost.error.type, so trace-based exporters (OpenTelemetry, Datadog) classify identically instead of re-deriving from the raw provider error.type. Pre-dispatch refusals produce no attempt span, so they reach the metrics but not the span.

Precision caveats on the inferred values

caller_model_unknown requires HTTP 404 on a request whose only addressable resource is the model: text completion, embedding, speech, transcription, image generation, rerank, count-tokens, and their streaming variants. Everything else classifies as caller_invalid_request, because its 404 may be about something other than the model:
  • file, batch, video and container operations address that resource directly;
  • OCR, image edit/variation and video generation reference a remote source asset (document_url, image.url, input_reference, video_uri);
  • chat completions and responses carry file ids, and responses additionally carry previous_response_id, which 404s on its own when stale.
Chat and responses are the notable absence. A wrong model name there lands in caller_invalid_request rather than being distinguished, because a 404 on those requests is genuinely ambiguous. Distinguishing it needs the provider to say so explicitly rather than Bifrost inferring it from a status code — see the per-provider normalization note above. The exact caller_model_not_available is unaffected: it is declared by Bifrost, not inferred, and covers every request type. The inferred values deliberately do not key off the provider’s own error.type / error.code. Those are passed through verbatim and disagree across providers for the same condition — OpenAI sends type="invalid_request_error" with code="model_not_found", Anthropic sends type="not_found_error" and has no code field at all, and Bedrock and Databricks each use their own vocabulary. OpenAI even sends invalid_request_error for a 401. Both fields remain available per-request on the span and in the logs. Two consequences, both accepted rather than papered over:
  • A provider that rejects an unknown model with a 400 lands in caller_invalid_request, not caller_model_unknown. This under-counts rather than mislabelling generic invalid-request traffic as a bad model name.
  • Some providers (Vertex) use 404 for “not found or your project does not have access to it”, so a permissions problem can land in caller_model_unknown. Separating the two needs per-provider error normalization; the Declared rows above do not depend on it.

Overhead Breakdown

bifrost_overhead_component_microseconds decomposes the same overhead measured by bifrost_overhead_latency_microseconds into per-component histograms. It carries the same base labels as bifrost_overhead_latency_microseconds (see Default Labels) plus an overhead_component label naming the internal component, and shares the same bucket boundaries. Summing every component for a given label set reconstructs the scalar total. overhead_component takes one of a fixed set of ten values. Each rolls the individual pipeline spans up into a category, matching the categories the Bifrost UI’s log-detail overhead breakdown groups into, so the metric and the UI agree: The set is bounded, so the list above is exhaustive. This metric is off by default. Enable it with the telemetry plugin’s overhead_breakdown_enabled config field (a sibling of metrics_enabled), or toggle Enable Overhead Breakdown on the pull-based tab of the Observability → Prometheus page in the UI.
The breakdown is computed from completed trace spans, so it only populates when tracing/observability is active for the request. With tracing off, bifrost_overhead_component_microseconds stays empty even when overhead_breakdown_enabled is on.

Bifrost MCP Metrics

Emitted for MCP (Model Context Protocol) tool calls executed through Bifrost: Labels: mcp_client (server label), mcp_tool_name, mcp_method (tools/call), error_type (auth_required / _OTHER on failure, empty on success), plus the governance labels virtual_key_id/virtual_key_name, team_id/team_name, customer_id/customer_name, business_unit_id/business_unit_name, project_id/project_name, and any custom labels. Only tool executions are recorded (lifecycle ping/list_tools and codemode tools are skipped); provider/model and network_transport are not labels here.

Default Labels

Most request-level Bifrost LLM metrics include these labels (the bifrost_key_rotation_events_total counter is an exception — see Key Rotation Events below for its narrower label set):
  • provider - LLM provider name
  • model - Model identifier
  • alias - Alias resolved to this model (empty if none)
  • method - Request type (chat, completion, embedding, etc.)
  • virtual_key_id / virtual_key_name - Virtual key identifiers
  • routing_engine_used - Comma-separated list of routing engines that contributed to the decision (e.g. governance, routing-rule, loadbalancing, model-catalog, core). core is emitted when the Bifrost orchestrator itself makes a routing decision — fallback transitions or retry transitions.
  • routing_rule_id / routing_rule_name - Routing rule that matched the request
  • complexity_tier - Complexity tier used for routing (SIMPLE / MEDIUM / COMPLEX); empty when no routing rule referenced complexity_tier
  • complexity_mechanism - How the effective complexity tier was determined (semantic, llm, session, or skipped when no tier was produced). The raw complexity score is deliberately not a label because it has unbounded cardinality; it remains available in request logs and trace attributes
  • selected_key_id / selected_key_name - API key that successfully served the request ("" when all attempts failed)
  • fallback_index - Fallback position
  • team_id / team_name - Team identifiers (empty when governance is not used)
  • customer_id / customer_name - Customer identifiers (empty when governance is not used)
  • project_id / project_name - Project the request was scoped to (empty when the request named no project). A request is scoped to at most one project, so these stay singular where team and customer identifiers can fan out
user_id / user_name are not included by default — see User Labels.
v1.5.0-prerelease4+: selected_key_id / selected_key_name are only populated when the request succeeds. On final errors both are empty — use the attempt_trail log field to see which keys were tried.

User Labels

user_id and user_name identify the end user a request was made on behalf of. They are available on every other observability surface — BigQuery columns, Splunk event fields, Datadog tags, span attributes — but are off by default on Prometheus metrics. Enable them with the telemetry plugin’s user_labels_enabled config field, or toggle User labels on the pull-based tab of the Observability → Prometheus page in the UI.
Values are populated by the enterprise auth middleware that resolves the calling user. On an OSS build with no user resolution, the labels are present but empty.

Key Rotation Events v1.5.0-prerelease4+

bifrost_key_rotation_events_total is incremented once per actual key rotation — i.e. when a per-key failure causes the next retry to switch to a different key. Rotation-triggering failures are bound to the specific key/account rather than the request:
  • 429 Too Many Requests — this key is rate-limited; another may have capacity.
  • 401 Unauthorized / 403 Forbidden — bad / revoked key, or key lacks permission.
  • 402 Payment Required — billing issue on this key’s account.
It is not incremented for:
  • terminal failures (no retry happens, including max_retries = 0 or every key permanently dead),
  • same-key retries on transient 5xx / network errors,
  • non-retryable request-bound 4xx (400/404/422/…).
Labels are attributed to the key that failed and triggered the rotation: To inspect every attempted key on a failed request (including terminal failures that did not rotate), read the attempt_trail field on the corresponding log entry instead. Example queries:

Push Gateway Setup

If you don’t have a Push Gateway running, deploy one:

Docker

Kubernetes (Helm)

Configure Prometheus to Scrape Push Gateway

Add to your prometheus.yml:
The honor_labels: true setting is important - it preserves the job and instance labels pushed by Bifrost instead of overwriting them with the Push Gateway’s labels.

Pull vs Push: When to Use Each

Why Push for Clusters?

When multiple Bifrost instances run behind a load balancer:
  1. Scraping randomness: Each scrape may hit different nodes, missing metrics from others
  2. Instance tracking: Push Gateway properly tracks per-instance metrics via instance label
  3. Aggregation: Downstream tools (Grafana, Datadog) can aggregate across all instances

Troubleshooting

Push Gateway Connection Failed

  • Verify the Push Gateway URL is correct and reachable from Bifrost
  • Check firewall rules between Bifrost and Push Gateway
  • Ensure Push Gateway is running: curl http://pushgateway:9091/metrics

Metrics Not Appearing

  • Verify the telemetry plugin is enabled (required for metrics collection)
  • Check Bifrost logs for push errors
  • Verify Prometheus is scraping the Push Gateway with honor_labels: true

Authentication Failed

  • Double-check username and password
  • Ensure basic auth is configured on the Push Gateway side
  • Check for special characters that may need escaping