Executive summary
21 months into the agentic era, Q2 left one question open: was the gain a one-time step or a slope? Q3 answers it. It is a slope, and it got steeper. Average output per employee rose 78% between June and September, the steepest three-month rise in the 21 months we have tracked. AI spend rose faster still, and routing is becoming a normal part of the stack.
Scroll sideways or zoom to read the full chart.
Three themes define the quarter:
1. The output curve steepened
Platform-wide, average monthly output per employee went from 76 units in June to 98 in July, 117 in August and 136 in September. July's 28% rise is the largest single-month gain in the series; August added another 20% and September 16%. Averaged over the quarter, output per employee ran 72% above the Q2 average. This is not just a change in who is on the platform. The constant-cohort median, which holds the set of organizations fixed, rose 3.1x over four quarters, faster than the blended average's 2.8x over the same period: teams themselves are accelerating. Part 2 has the detail.
2. Spend outran output, but cost growth is slowing
Average AI spend per employee rose from $257 a month in June to $637 in September, 2.5x in three months, while output rose 1.8x. The rate of increase is what changed. Spend per employee rose 69% in July, 34% in August, and 9% in September, while output per employee rose 28%, 20%, and 16%. September is the first month this quarter in which output grew faster than spend. Cost growth is slowing, not yet bending: one month is not a turn. Of each $100 of agent session spend, $23 ships cleanly, against $26 in Q2 (Q3 excludes one outlier organization).
3. Routing is becoming normalized
In September, the Weave Router handled about 1.62 million requests across 52 installations and 234,000 agent sessions. It served 82% of them with a different model than the one requested, and routed traffic billed about 44% less than the requested models would have. 53% of input tokens were prompt-cache reads, which is why Router 2.0, released in September, pins sessions to protect warm caches. On Terminal-Bench 4.0 it solved 62.1% of tasks against 60.6% for GPT-6 Astra, at 52% of the cost and 2.2x the speed.
The bottom line for leaders
Q2's conclusion holds: the gain is real, it is table stakes, and what separates organizations is conversion. Q3 adds urgency. The curve is steepening, so the distance between organizations that measure value and those that measure volume is widening each month. The bill is growing faster than the output, so the case for governing spend is now a budget question rather than a principle.
Read the full report
Enter your work email to unlock every chart, benchmark and methodology note in the report.