When we tell people Legwork cut our reasoning model's input tokens by 98.6%, the first question is usually "on what?" Fair question. Here's the short version, and the parts we'd rather you hear from us than discover yourself.
The setup
On 5 October we took a real working corpus: a 2,275-file Vietnamese novel manuscript and its notes, about 35 MB of text. The author gave us 10 questions they actually ask day to day. Each question ran twice in two modes. In one, the reasoning model did all the searching and reading itself. In the other, it handed the digging to a Legwork worker and worked from the report. That's 40 runs, each with a 10-minute budget, and the author reviewed all 40 answers blind.
Searches and reads every file itself
Writes the answer
Hands off the digging
Searches, reads, cross-checks
Short, with sources
Writes the answer
What moved
Across 20 runs per mode, the reasoning model's input went from 37.45M tokens to 0.51M, and its requests from 776 to 68. Even counting only tokens that weren't cache hits, input dropped 88.4%. The average review score was identical, 4.12 out of 5 in both modes, and the delegated answer was preferred in 6 of 10 questions.
What it costs
We priced the measured tokens at list API rates as of 8 October 2026. The reasoning model was Qwen3.8-Max ($2 per 1M input tokens, $0.25 cached, $6 output, Alibaba Cloud Singapore). The worker was Agnes 3.0 Flash ($0.05 input, $0.005 cached, $0.15 output).
Doing all the digging itself, Qwen3.8-Max cost $17.19 across 20 runs, about $0.86 per question. With Legwork, its share fell to $1.40. The worker read 43.1M tokens of prompt to do the digging, mostly cache hits, wrote 0.35M, and added $0.61. That's $2.01 in total, about $0.10 per question, or 88% less.
| At list API rates | Reasoning model | Worker | Total | Per question |
|---|---|---|---|---|
| Does it all | $17.19 | $0 | $17.19 | $0.86 |
| With Legwork | $1.40 | $0.61 | $2.01 | $0.10 |
Two honest footnotes. Money falls less than tokens (88% versus 98.6%) because most of the "does it all" input was cache hits, which are cheap. And the "does it all" figure is a floor: 16 of its 20 runs hit the 10-minute budget before finishing.
What didn't move
It wasn't faster: both modes took about the same total time. The work didn't vanish either. The worker read about 8.2 MB and returned 0.68 MB of reports, so the reading still happened, just on a low-cost model you choose instead of your main one.
Where it was worse
Two of the 20 delegated answers contained an inaccuracy; none of the "does it all" answers did. That's the reason every Legwork finding carries an exact source. You can check a claim in seconds, and re-running the worker is cheap.
What we haven't tested yet
This was one corpus with one reviewer, and the reasoning model was driven through its API, not inside the Claude or ChatGPT apps. Testing inside the apps is in progress, and we'll publish those numbers here when we have them.
If you want every table, the full methodology is public.