Legwork benchmark
How much work moves off your main model when it delegates the digging, and what happens to answer quality. Internal test, run 5 October 2026, reviewed 6 October 2026.
Real questions, real documents.
Documents
A 2,275-file Vietnamese-language novel manuscript and its notes, about 35 MB of text.
Questions
10 questions the author actually asks day to day about the manuscript.
Two modes
The reasoning model does all the searching and reading itself, or it delegates the digging to a Legwork worker and works from the report.
Runs
Each question twice in each mode: 40 runs. Each run had a 10-minute budget.
Quality review
All 40 answers reviewed blind by the manuscript's author, who knows the material best.
Results
Totals over 20 runs per mode.
| Reasoning model | Does it all | With Legwork | Change |
|---|---|---|---|
| Requests | 776 | 68 | −91% |
| Input tokens | 37,452,931 | 510,012 | −98.6% |
| Input tokens, not cached | 3,060,099 | 354,364 | −88.4% |
| Output tokens | 411,559 | 107,971 | −74% |
| Tool results read into its context | 4.35 MB | 0.68 MB | −84% |
| Total time | 12,827 s | 12,733 s | about the same |
In the "does it all" mode, 16 of 20 runs were still investigating when the 10-minute budget ran out, so its figures are a lower bound. Most of its input was cache hits, which providers usually bill at a lower rate; the not-cached row shows that view.
Same average, more variation.
| Blind review | Does it all | With Legwork |
|---|---|---|
| Average score (1–5) | 4.12 | 4.12 |
| Questions where this answer was preferred | 4 of 10 | 6 of 10 |
| Answers with an inaccuracy | 0 of 20 | 2 of 20 |
| Answers with a minor omission | 1 of 20 | 2 of 20 |
One question had a best answer in each mode and is counted for both; one had no clear favourite. Every question got a usable answer in both modes. Both inaccuracies were in delegated runs, which is why Legwork reports carry a source for every finding and make re-running the worker cheap.
What this test doesn't show.
- Not measured inside the Claude or ChatGPT apps. The reasoning model was driven through its API. Testing inside the apps is in progress.
- Not a speed gain. Both modes used about the same time.
- The work moves; it doesn't vanish. The worker read about 8.2 MB and returned 0.68 MB of reports. That work lands on a low-cost model you choose instead of your main one.
- Worker cost depends on your choice of model. We don't name models here because you pick your own, and results will vary with that choice.
- One corpus, one reviewer. The questions were hard, real ones. Everyday questions may be lighter; we haven't measured that yet.