Skip to content

Repository files navigation

BitRouter

Build status Crates.io License: Apache-2.0 Twitter Discord Hugging Face Docs LinkedIn Book a call

An open-source, context-aware model router that learns and adapts to your agent workflows.

You're tokenmaxxing in production. Every step of every loop bills at frontier prices — file reads, tool calls, sub-agent hops, retries. Most don't need it. BitRouter routes each call, tool, and agent to the cheapest path that still reaches the goal, and tightens that routing as the loop runs.

Works with any harness, any model, any loop. Cost is live today — latency and accuracy are next.

One gateway for everything the loop consumes

BitRouter routes model calls. But which model a call should get depends on where the loop is — and the loop step is the last tool, skill, or sub-agent it touched. So the same gateway that governs those is what gives the router its context. Other routers see only the first of these three:

  • Models — route LLM calls across providers, accounts, and wire protocols: OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and Google Gemini. (the classic router, cross-protocol — any request format to any upstream, and back)
  • Capabilities — an MCP gateway that carries both tools and skills: skills ride the MCP skills extension (SEP-2640) over MCP Resources, so tools and skills are one governed, routable namespace instead of hardcoded endpoints. (Skills transport is unstable while SEP-2640 is in review.)
  • Agents — an ACP gateway: sub-agents become first-class routable primitives, so a task can go to the sub-agent that best fits the loop's objective — just as a call routes to the best-fit model. (Local sub-agents over stdio today; remote gateways arrive with ACP v2.)

Optimizing a loop isn't just model selection — it's choosing the model, the tool, and the sub-agent that best serve the loop's objective at every step that gets it to its goal.

The self-improving loop

BitRouter wraps your agentic loop in a second loop. bitrouter.yaml declares providers, presets, and whether the process may publish; policy-lock.yaml is the only live route authority. The routing key is context-aware and lives as code: it is the step in the loop, not just the model name.

policy:
  path: ./policy-lock.yaml
  mode: adaptive                 # authorizes explicit publication only
presets:
  auto:
    model: openai-codex:gpt-5.6-sol
    policy: auto

The v3 lock behind bitrouter/auto contains the tier targets, canonical agent_trace routes, capability guardrails, and a decision certificate for every explicit route. A target may be a scalar model or an exact (model, effort) pair; bitrouter/auto:cost selects the cost variant when one is defined, while explicit physical model IDs remain passthrough.

Against that spec BitRouter provides the control plane for an act → observe → evaluate → improve cycle:

  • Act — the router reads the lock and rewrites each @preset[:variant] call to its tier's model: policy routing, cross-protocol translation, multi-account failover.
  • Observe — telemetry attributes every hop with cost, tokens, latency, and outcome, exported to Prometheus or any OTLP backend.
  • Evaluate — the generic eval exchange lets task-native tests, humans, enterprise systems, or an external agentic judge submit the same versioned outcome contract. BitRouter admits, disputes, and snapshots evidence; it does not pretend one bundled judge is universal.
  • Improve — the policy compiler turns a frozen admitted-evidence snapshot into a deterministic, certificate-backed policy-lock.yaml candidate (an npm-style manifest/lock split, git-owned). Review and publication are explicit; the evidence database never changes a live route.

You choose what the external evaluator measures — cost, latency, quality, or a private objective — while the active lock remains the only authority for live policy routing. BitRouter adapts by proposing a new lock, never by mutating the one in force: the live route never changes implicitly between publications, and every change is a diff you can read and revert.

Benchmarks

Today cost is the validated objective: on Terminal-Bench 2.1, gpt-5.5 with BitRouter cut cost 32.8% at near-parity accuracy (−1.1 pp), by offloading routine steps to a cheaper model. Latency and accuracy objectives — and more base models — are landing next.

Base model Cost vs baseline Latency vs baseline Accuracy vs baseline
gpt-5.5 **−32.8%**¹ coming soon coming soon
gpt-5.6 coming soon coming soon coming soon
claude-opus-5 coming soon coming soon coming soon
claude-sonnet-5 coming soon coming soon coming soon
claude-fable-5 coming soon coming soon coming soon

¹ Cost-optimization run on Terminal-Bench 2.1: −32.8% zero-cache imputed cost (audited range 28.6–32.8% by cache share) at near-parity accuracy, −1.1 pp (76.1% vs 77.3%, within single-attempt noise).

This is a mechanism study under a modified protocol, not a Terminal-Bench leaderboard submission — read the experiment limitations before citing the numbers. Full reports live in benchmarks/; complete traces, tool calls, usage, policy decisions, configs, and checksums are in the BitRouterAI/benchmarks dataset.

Comparison

Automatic model selection is no longer the differentiator — every router below picks a model for you. What separates them is what the decision reads and whose data makes it better. Almost all of them classify the prompt. BitRouter routes on the loop step — where the agent is in its trajectory, keyed by the last tool it called — and improves from evaluations you admit, into a lock file you own.

BitRouter OpenRouter Auto Not Diamond Code vLLM Semantic Router LiteLLM Auto
Routing signal The loop step — last tool called, over the agent trace Prompt classified into ~30 task types Session state, token counts, task complexity, KV-cache state Prompt signals: classifiers, embeddings, keyword and metadata rules Prompt embedding similarity to labeled example routes
Improves from Your admitted evaluation outcomes, per loop The community's last-7-days spend across all of OpenRouter Your implicit accept/reject feedback, in a hosted model Router models you train offline and redeploy Labeled examples plus one fitted score threshold
Optimizes Any objective you submit (cost validated today) Cost tier vs. capability Cost and quality jointly Quality, cost, latency, privacy, safety Cost vs. quality at a chosen threshold
Policy you own Git-owned policy-lock.yaml — readable, diffable, explicit publish Vendor-side, tunable by knobs Vendor-side Config recipes plus trained artifacts Proxy config plus fitted threshold

OpenRouter Auto and Not Diamond are hosted services. BitRouter, vLLM Semantic Router, and LiteLLM are open-source and self-hostable; BitRouter is Rust.

Install

# macOS / Linux
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/bitrouter/bitrouter/releases/latest/download/bitrouter-installer.sh | sh

# Homebrew
brew install bitrouter/tap/bitrouter

# npm
npm install -g bitrouter
From source (Cargo)
cargo install bitrouter

Quick Start

BitRouter is a local proxy between your agent and every LLM provider. One env-var swap — no harness changes required:

- OPENAI_BASE_URL=https://api.openai.com/v1   # hardwired to one provider, no fallback
+ OPENAI_BASE_URL=http://localhost:4356/v1    # all providers, automatic failover

CLI

BitRouter runs as a local daemon — start it with your own keys or a Cloud sign-in.

Bring your own keys (BYOK) — auto-detected from the environment, no config file needed:

export OPENAI_API_KEY=sk-...    # ANTHROPIC_API_KEY / GEMINI_API_KEY also work
bitrouter start                 # proxy running at http://localhost:4356

Or sign in to BitRouter Cloud — use browser OAuth interactively or store an existing API key in CI:

bitrouter cloud login           # RFC 8628 device flow against api.bitrouter.ai
bitrouter cloud login --api-key "$BITROUTER_API_KEY"  # non-interactive CI login
bitrouter start                 # `bitrouter` provider auto-enables once signed in

The same credential also drives a gh api-style raw client—no daemon required:

bitrouter cloud api /v1/models
bitrouter cloud api /v1/chat/completions --input request.json

Point your agent runtime at http://localhost:4356 and any available provider is live. For advanced routing rules, guardrails, or multi-account failover, scaffold a config with bitrouter init (writes ./bitrouter.yaml).

bitrouter start / stop / restart        # daemon lifecycle
bitrouter status --watch                # live request stream + spend
bitrouter route <model>                 # trace how a model name resolves
bitrouter key sign --user <id>          # mint a scoped brvk_ API key
bitrouter cloud keys list               # manage API keys
bitrouter cloud usage                   # inspect spend and tokens
bitrouter cloud billing balance         # check credits
bitrouter cloud api /v1/models          # call Cloud APIs directly

See docs/CLI.md for the full command reference, flags, and config resolution.

Agent Skill

BitRouter ships an Agent Skill/bitrouter — so AI coding agents can install, configure, migrate to, and troubleshoot BitRouter on their own. It lives in this repo at skills/bitrouter/, kept in sync with the code.

npx skills add bitrouter/bitrouter    # via the generic skills CLI
# ...or add this repo as a plugin marketplace in Claude Code / Codex

MCP

Use BitRouter from any MCP client — it exposes complete, list_models, and status as MCP tools (the origin server, distinct from the MCP gateway that proxies your own MCP servers):

bitrouter mcp serve                    # stdio → local daemon at 127.0.0.1:4356
bitrouter mcp install --client claude  # print the Claude/Cursor mcpServers config block

Add --transport http to target the multi-tenant cloud backend.

API

BitRouter exposes an OpenAI- and Anthropic-compatible HTTP API on http://localhost:4356, so any SDK or client works unchanged. The full endpoint reference and OpenAPI spec live in bitrouter/bitrouter-docs (rendered at bitrouter.ai).

Workflow templates

Ready-made policy specs for common agentic workflows start in templates/auto-router/: a predictive bitrouter/auto / bitrouter/auto:cost ladder using GPT-5.6 as the strong tier, Kimi K3 as balanced, and DeepSeek V4 Pro as economy. Treat it as a starting point and evaluate it against your own loop before publishing a live policy.

Models & providers

BitRouter routes to a model, not a provider. Each model below is served by many providers — its own lab, hyperscalers (AWS Bedrock, Alibaba Cloud), gateways (OpenRouter, OpenCode), and serverless clouds — and BitRouter picks the cheapest route per call. The bars are those routes: the more providers serve a model, the more room the router has to move.

BitRouter model catalog grouped by lab; bar length is the number of providers routing each model

Bring your own key to any of them, or use one BitRouter Cloud account with no keys at all. Frontier models from OpenAI, Anthropic, Google, and xAI also route over a subscription sign-in (Claude Pro/Max, GitHub Copilot, ChatGPT Codex) instead of a key.

The chart is generated from dist/registry/ on every catalog change — it is never hand-maintained, so it cannot drift from what the router actually resolves. Full catalog in registry/.

Harness integrations

BitRouter runs under Claude Code, Codex, and the rest — not instead of them. Any agent runtime that speaks OpenAI or Anthropic APIs works with it out of the box — set OPENAI_BASE_URL=http://localhost:4356/v1 and you're done. For the four harnesses below, bitrouter launch does the wiring for you: it starts the harness's own native TUI with its traffic already pointed at the daemon, and never edits the harness's config files.

Name Description
Claude Code Child env overrides (ANTHROPIC_BASE_URL) — see the LLM gateway guide for the manual form
OpenAI Codex One-shot -c overrides — see custom model providers for the manual form
OpenCode Synthesized OPENCODE_CONFIG; models via models.dev
Pi-Agent Synthesized PI_CODING_AGENT_DIR — see the model configuration guide for the manual form

launch supports these four because every promise it makes — routing and gateway injection — has to be re-verified per harness against upstream releases nobody controls. Four is a surface that stays honest.

Headless ACP sub-agents use bitrouter spawn instead. The full provider and harness catalog lives in github.com/bitrouter/bitrouter/registry.

Features

Beyond the gateways above, the production controls for running agents unattended:

  • Multi-account failover + load-balancing — reroute mid-run; a rate-limit at file 140 never re-pays for files 1–139
  • Virtual keys (brvk_) scoped per agent or user — no agent holds an upstream key
  • Per-agent spend caps + loop guards to contain runaway cost
  • Injection + output guardrails at the router, before requests leave your network
  • Zero-config auto-detection + custom OpenAI-/Anthropic-compatible providers

Talk to founders

Try BitRouter Cloud → or reach out directly:

Want a first-party provider integration, or building an open-source agent/harness? Email kelsenliu@bitrouter.ai or book a meeting — open-source builders get up to 50% off for you and your community.

Development

  • docs/DEVELOPMENT.md — workspace architecture and SDK internals
  • CONTRIBUTING.md — contribution workflow, issue reporting, and provider updates
  • CLAUDE.md — guidance for AI coding agents working in this repository
  • skills/ — the /bitrouter Agent Skill (source of truth)

Star History

Star History Chart

License

Licensed under the Apache License 2.0.

About

The context-aware model router that learns & adapts to your long-horizon coding agent workflows. Lightweight & extensible, works with any harnesses, any models, any loops.

Topics

Resources

Contributing

Stars

221 stars

Watchers

1 watching

Forks

Releases

Used by

Contributors

Languages