Weave Router vs Not Diamond: which model router is right for your team?
Both are learned routers that send each prompt to the model most likely to deliver. Not Diamond is an API you integrate into your own app or agent. Weave Router works inside Codex, Claude Code, and Cursor with no integration at all. The right choice depends on where your tokens burn. Here’s the full breakdown.
Weave Router routes each prompt to the best quality-per-token model right inside Codex, Claude Code, and Cursor. Not Diamond is a learned routing API, priced at $0.05 per million tokens routed, that you build into your own application or agent.
Teams typically choose Weave Router when the tokens they want back are burned by engineers in coding tools. Teams typically choose Not Diamond when they’re building their own AI product and want routing as an API inside it.
At a glance
How Weave Router and Not Diamond compare
| Compare | Weave Router | Not Diamond (notdiamond.ai) |
|---|---|---|
| Core approach | Routes every prompt to the best quality-per-token model | Learned router API that picks the LLM most likely to perform per query |
| Where it runs | Inside Codex, Claude Code, and Cursor, no code changes | An API you integrate into your own app or agent |
| Integration effort | None: 5-minute setup, no SDK | Code integration; custom routers trained on your prompts and eval scores |
| Model coverage | Frontier and efficient models across major providers | 60+ models across providers |
| Pricing | Free to get started | $0.05 per million tokens routed |
| Savings story | Roughly doubles how far a token budget goes on coding work | Reports 20–40% inference savings for coding agents |
| Source availability | Source code available | Proprietary hosted API |
| Best for | Engineering teams using AI coding assistants and agents | Teams building their own AI products and agents |
What is Weave Router?
Weave Router is the prompt router that doubles your token runway. It routes each prompt to the best quality-per-token model, right inside the tools your engineers already use: Codex, Claude Code, and Cursor. There’s no API to integrate and no eval dataset to assemble: easy prompts stop paying frontier prices from the moment it’s installed. Setup takes about five minutes, the source code is available, and it’s free to get started. Weave Router is built by Weave, the engineering intelligence platform trusted by 500+ organizations.
What is Not Diamond?
Not Diamond (notdiamond.ai) is a learned model router offered as an API. You call it from your own application or agent and it returns the model most likely to perform for that query, across 60+ models, for $0.05 per million tokens routed. To get the most from it you can train custom routers on your own prompts, candidate responses, and evaluation scores, and it reports 20–40% inference savings for coding agents. It’s a strong choice when you’re building your own AI product. But if the tokens you want back are burned inside Codex, Claude Code, and Cursor, there’s nothing to integrate an API into.
The key difference
An API for builders vs. a router for engineering teams
Not Diamond’s approach
Not Diamond gives developers a routing primitive. You integrate the API into your product, optionally train a custom router on your prompts and eval scores, and every call gets matched to the model most likely to perform. Powerful, but it assumes there’s a codebase you control sitting between your users and the models.
Weave Router’s approach
Your engineers’ coding tools aren’t your codebase; you can’t integrate an API into Cursor. Weave Router routes those prompts anyway: it drops into Codex, Claude Code, and Cursor, scores every prompt, and sends it to the best quality-per-token model, working out of the box with no eval datasets to assemble.
If you’re building an AI product, integrate a router. If your engineers are burning tokens in coding tools, install one.
Why Weave Router
Why teams choose Weave Router over Not Diamond
Nothing to integrate
Not Diamond assumes you control the code calling the models. Weave Router works where you don’t: inside Codex, Claude Code, and Cursor.
No eval datasets required
Custom Not Diamond routers are trained on prompts, candidate responses, and eval scores you assemble. Weave Router is tuned on engineering work out of the box.
Doubles your token runway
Quality-per-token routing on real coding workloads takes the same budget roughly twice as far, beyond the 20–40% Not Diamond reports for coding agents.
No per-token routing fee
Not Diamond charges $0.05 per million tokens routed. Weave Router is free to get started, with no meter running on the routing itself.
Source code available
Not Diamond is a hosted API. Weave Router’s source is available, so your team can inspect exactly how routing decisions are made.
5-minute setup
Sign up at router.workweave.ai and your engineers’ prompts are being routed the same afternoon. No sprint required.
When Not Diamond might be the better fit
We’d rather you pick the right tool than just pick us. If you’re building your own AI product or agent and want a learned router inside it, especially if you have the prompts and eval scores to train a custom router, Not Diamond is a strong, focused choice with low per-call overhead and wide model coverage. Weave Router doesn’t compete for that job. But your engineers’ coding tools aren’t code you control, and their tokens are often the biggest line item. Routing those prompts is exactly what Weave Router was built for. Plenty of teams use both.
Frequently asked questions
Both are learned routers, but they live in different places. Not Diamond is an API you integrate into your own application or agent. Weave Router installs into Codex, Claude Code, and Cursor, so it can route the coding prompts no API integration can reach.
No. There’s no SDK and no API to call. Weave Router sets up inside Codex, Claude Code, and Cursor in about five minutes.
Not with Weave Router. It’s tuned on engineering work, code generation, refactors, reviews, and agent runs, and routes well from the first prompt. Not Diamond’s custom routers improve with prompts, candidate responses, and eval scores you supply.
Weave Router routes each prompt to the best quality-per-token model, which roughly doubles how far a token budget goes on typical engineering workloads. Try the savings calculator on the Router product page for an estimate based on your usage.
Yes. They route different traffic: Not Diamond routes calls inside the products you build, while Weave Router routes your engineers’ prompts inside their coding tools.
Source code available. Start taking your token budget twice as far.
Get started in 5 minutes or book a demo with our team.