← Legwork

Legwork benchmark

How much work moves off your main model when it delegates the digging, and what happens to answer quality. Internal test, run 5 October 2026, reviewed 6 October 2026.

What we tested

Real questions, real documents.

Documents

A 2,275-file Vietnamese-language novel manuscript and its notes, about 35 MB of text.

Questions

10 questions the author actually asks day to day about the manuscript.

Two modes

The reasoning model does all the searching and reading itself, or it delegates the digging to a Legwork worker and works from the report.

Runs

Each question twice in each mode: 40 runs. Each run had a 10-minute budget.

Quality review

All 40 answers reviewed blind by the manuscript's author, who knows the material best.

Reasoning-model workload

Results

Totals over 20 runs per mode.

Input tokens

−98.6%
Does it all
37.45M
With Legwork
0.51M

Requests

−91%
Does it all
776
With Legwork
68
Same quality

Equal average score in a blind review. The delegated answer was preferred in 6 of 10 questions.

Reasoning modelDoes it allWith LegworkChange
Requests77668−91%
Input tokens37,452,931510,012−98.6%
Input tokens, not cached3,060,099354,364−88.4%
Output tokens411,559107,971−74%
Tool results read into its context4.35 MB0.68 MB−84%
Total time12,827 s12,733 sabout the same

In the "does it all" mode, 16 of 20 runs were still investigating when the 10-minute budget ran out, so its figures are a lower bound. Most of its input was cache hits, which providers usually bill at a lower rate; the not-cached row shows that view.

Answer quality

Same average, more variation.

Blind reviewDoes it allWith Legwork
Average score (1–5)4.124.12
Questions where this answer was preferred4 of 106 of 10
Answers with an inaccuracy0 of 202 of 20
Answers with a minor omission1 of 202 of 20

One question had a best answer in each mode and is counted for both; one had no clear favourite. Every question got a usable answer in both modes. Both inaccuracies were in delegated runs, which is why Legwork reports carry a source for every finding and make re-running the worker cheap.

Limits

What this test doesn't show.

  • Not measured inside the Claude or ChatGPT apps. The reasoning model was driven through its API. Testing inside the apps is in progress.
  • Not a speed gain. Both modes used about the same time.
  • The work moves; it doesn't vanish. The worker read about 8.2 MB and returned 0.68 MB of reports. That work lands on a low-cost model you choose instead of your main one.
  • Worker cost depends on your choice of model. We don't name models here because you pick your own, and results will vary with that choice.
  • One corpus, one reviewer. The questions were hard, real ones. Everyday questions may be lighter; we haven't measured that yet.