- NPX
- Docker
2.2.4
✨ Features
- Anthropic Between-Tools Thinking -
reasoning.type: "between_tools"on chat and Responses requests, andthinking: {"type": "between_tools"}on the Anthropic drop-in route, are forwarded to Anthropic, Bedrock and Vertex with the caller’s effort passed independently. Models without it getdisabledor no thinking field, so a fallback to an older model never fails. The datasheetsupports_between_tools_thinkingfield can override this (#7665) - Tool Search for GPT-5.4 and Later -
defer_loadingon function and MCP tools now reaches OpenAI, Azure, Bedrock and Bedrock Mantle for gpt-5.4, gpt-5.5, gpt-5.6 and gpt-6 models, so the model can search deferred tools instead of loading every tool eagerly. Other OpenAI-compatible backends still have it stripped. The datasheetsupports_tool_searchfield can override this (#7534) - Skipped Routing Fallbacks Are Logged - A rule fallback that names no known provider is now reported in the request’s routing log with the rule name and the configured entry, instead of being skipped silently (#7546)
🐞 Fixed
- Routing Fallbacks Dropped After Restart - Legacy
provider/modelfallback strings are re-parsed at route time, so a custom provider registered after routing rules were decoded at boot is no longer skipped while the API still listed the fallback. Object-form fallbacks naming an unregistered or blank provider are now rejected on create and update (#7543, #7544, #7545) - Bedrock Thinking Tokens - Requests on
/bedrock/model/{id}/invokeand its streaming sibling with extended thinking on are served through InvokeModel, andusage.output_tokens_details.thinking_tokensis reported in unary responses and in themessage_startandmessage_deltaevents (#7691) - Encrypted Reasoning Retry After a Provider Switch - The strip-and-retry for replayed reasoning now fires on any 400 that names a reasoning token, instead of matching each provider’s wording, so a mid-conversation switch such as bedrock to bedrock_mantle heals instead of returning the 400 to the client. Gemini and Vertex thought signatures carried inside call ids are stripped as well (#7680)
- OpenRouter Error Messages - The upstream provider’s own error message is lifted out of
error.metadata.raw, replacing the generic “Provider returned error” (#7680) - Truncated Turns on Anthropic and Gemini - Responses turns cut off by
max_output_tokensor a refusal now report statusincompletewithincomplete_details, streams end withresponse.incomplete, and the Bedrock, Gemini and Cursor drop-in routes translate that into their own stop reason. A Gemini stream that ends without a finish reason no longer reads as a clean stop (#7677) - File Data on OpenAI-Compatible Providers -
file_datasent as bare base64 with afile_typeis folded into adata:URL, which OpenAI and Databricks require, instead of being rejected with “Invalid base64 data URL format” (#7682) - Gemini Histories Replayed to OpenAI - A
function_callinput item whose id does not start withfcno longer fails OpenAI validation natively or through a fallback. The id is dropped andcall_idis kept so outputs still pair (#7676) - Gemini Thought Signatures on Images and Files - A thought signature on an inline image or file part stays on that content block and round-trips back to Gemini, instead of being dropped or emitted as a separate reasoning item (#7692)
- Governance Resets During Startup - Startup resets and the periodic reset worker now run only after governance state is fully hydrated, and team-owned budgets and rate limits keep the team’s calendar alignment after a restart, so calendar-aligned limits are no longer reset on a creation-anchored boundary (#7615, #7637)
- Model Histogram Unnamed Series - Rows without a model, such as list_models, file and batch operations, are excluded from the model histogram (#7632)
- Ungoverned Virtual Key Creation - A user without an access profile no longer sees locked governance fields when creating a virtual key. The form locks only when a profile actually governs, and a failed policy lookup shows a warning with a retry instead of locking (#7413)
🗄️ Database Migrations
- No new database migrations in this release.
🐙 Closed GitHub Issues
1.11.0
- feat: reasoning.type “between_tools” on chat and Responses requests, forwarded as Anthropic thinking.type on Anthropic and Bedrock with the caller’s effort passed independently, downgraded to disabled or omitted on models without it, with a SupportsBetweenToolsThinking datasheet override (#7665)
- feat: Bedrock invoke ingress serves Anthropic thinking requests through InvokeModel and reports usage.output_tokens_details.thinking_tokens in unary and streaming responses (#7691)
- fix: the encrypted reasoning fail-soft retry fires on any 400 that names a reasoning token by family word instead of per-provider verdict phrasing, and strips Gemini and Vertex thought signatures carried in ts call ids (#7680)
- fix: OpenRouter errors surface the upstream provider’s own message from error.metadata.raw instead of the generic “Provider returned error” (#7680)
- fix: Anthropic and Gemini Responses turns cut off by max_output_tokens or a refusal report status incomplete with incomplete_details, terminate streams with response.incomplete and mark the last output item incomplete; Gemini streams ending without a finish reason no longer read as a clean stop (#7677)
- fix: file_data sent as bare base64 with a file_type is folded into a data URL for OpenAI-shaped providers, which rejected the bare payload (#7682)
- fix: OpenAI Responses drops a function_call input item id that does not begin with fc so Gemini streaming histories replay cleanly; call_id is kept (#7676)
- fix: Gemini thought signatures on inline image and file parts stay on the content block and round-trip back to Gemini (#7692)
1.7.5
- fix: routing fallbacks in the legacy provider/model string form are re-parsed at route time, so custom providers registered after the rules were decoded at boot are no longer dropped; object-form fields are trimmed on decode (#7543)
- fix: model histogram queries exclude rows with an empty model, so list_models, file and batch operations no longer appear as an unnamed series (#7632)
- chore: upgraded core to v1.11.0
0.3.4
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.8.4
- fix: startup resets and the periodic reset worker start through StartResetWorkers after all governance state is hydrated, so a calendar-aligned budget is never reset on a creation-anchored boundary during boot (#7637)
- fix: team-owned budgets and rate limits keep the team’s calendar_aligned value after a restart or config reload (#7615)
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.6.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.8.4
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.7.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.6.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.1.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.5.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.1.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.1.4
- fix: rule fallbacks resolve legacy provider/model strings at route time, so custom-provider fallbacks survive a restart (#7543)
- feat: a fallback that names no known provider is reported in the request’s routing log with the rule name and the configured entry instead of being skipped silently (#7546)
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.6.7
- chore: upgraded core to v1.11.0 and framework to v1.7.5
1.8.3
- chore: upgraded core to v1.11.0 and framework to v1.7.5

