Intelligent
model router for

coding agents

Unlock faster cycles, at half the cost with frontier quality

  • 51% of the cost
  • 2.3x Faster
  • Astra Level Quality
Loading router example…refactor the auth module: DeepSeek, score 91, cost $$. debug a flaky integration test: Claude, score 96, cost $$$. summarize this pull request: Llama, score 74, cost $. triage 40 support tickets: Gemini, score 82, cost $. migrate the config to YAML: Kimi, score 79, cost $$. draft the release notes: GPT-5, score 88, cost $. explain this stack trace: Claude, score 90, cost $$.
Weave

Weave reads every turn your agent takes and sends it to the cheapest model that’ll get it right — saving the frontier ones for when it counts.

The receipt

One task apart. Half the bill.

Same Codex harness, same tasks, two attempts each — Router and GPT-6 Astra paired on the same day and package.

Compared with GPT-6 Astra

48%less cost per trial than AstraOne task apart on 66. Router solved 62.1%; Astra solved 60.6%.
Cost per trialLower is better
  1. WeaveRouter$5.22
  2. GPT-6 Astra$10.03
  3. OpenRouter Auto$5.52

Terminal-Bench 4.0

54%less cost per trial than Astra$2.31 per trial, compared with Astra’s $5.04.
Cost per trialLower is better
  1. WeaveRouter$2.31
  2. GPT-6 Astra$5.04
  3. OpenRouter Auto$1.49

SWE-Atlas

Cost and time are per trial. Quality is pass@2: a task counts if either of two attempts passed. Router and Astra are paired; OpenRouter is a separate run on the same tasks and graders. Charts share a scale within each metric.

Reproduce it

How it works

Most routers answer one question. This one answers four.

From the blogWeave Router 2.0: Astra quality at half the cost
  • How hard is this turn?

    A complexity classifier trained on 10× the sessions of 1.0 scores it in single-digit milliseconds.

  • Will switching cost more than it saves?

    Cache-aware. It only routes down when the saving beats the cost of rebuilding the cache.

  • Is the cheap model actually finishing?

    An escalation classifier watches the task. Stall, loop, or miss → back to a frontier model.

  • Which subscription has quota?

    Claude inside Codex. GPT inside Claude Code. It drains the flat-rate seats you already own first.

Integration

Your tools stay the same.

One command detects Claude Code, Codex, and Cursor and configures each. Add the keys or subscriptions you want to route across.

npx @workweave/router

Estimate

Now put your traffic through it.

An estimate against the current pool.

Coding harness

Model mix today

Monthly token spend

$12,000
$2k$250k
$6,360/mo

53% lower monthly spend · $76,320 annualized

Current
$12,000
Routed
$5,640

74% of requests routed to the pool

Estimate from routed-traffic benchmarks across the current pool. Actual savings depend on workload, prompt mix, and provider pricing.

Astra quality at half the cost.

One command to start, 5% of routed spend, Elastic License 2.0 to self-host. Teams of 50+ get an FDE to set it up with them.