Intelligent
model router for
coding agents
Unlock faster cycles, at half the cost with frontier quality
- 51% of the cost
- 2.3x Faster
- Astra Level Quality
Weave reads every turn your agent takes and sends it to the cheapest model that’ll get it right — saving the frontier ones for when it counts.
The receipt
One task apart. Half the bill.
Same Codex harness, same tasks, two attempts each — Router and GPT-6 Astra paired on the same day and package.
SWE-Atlas
How it works
Most routers answer one question. This one answers four.
How hard is this turn?
A complexity classifier trained on 10× the sessions of 1.0 scores it in single-digit milliseconds.
Will switching cost more than it saves?
Cache-aware. It only routes down when the saving beats the cost of rebuilding the cache.
Is the cheap model actually finishing?
An escalation classifier watches the task. Stall, loop, or miss → back to a frontier model.
Which subscription has quota?
Claude inside Codex. GPT inside Claude Code. It drains the flat-rate seats you already own first.
Integration
Your tools stay the same.
One command detects Claude Code, Codex, and Cursor and configures each. Add the keys or subscriptions you want to route across.
npx @workweave/router- claude code
- codex
- cursor
- Claude
- GPT
- DeepSeek
- GLM
- Gemini
Estimate
Now put your traffic through it.
An estimate against the current pool.
Coding harness
Model mix today
Monthly token spend
53% lower monthly spend · $76,320 annualized
74% of requests routed to the pool
Estimate from routed-traffic benchmarks across the current pool. Actual savings depend on workload, prompt mix, and provider pricing.


