Lineage artwork for Proof: which workers did the best work, and what it cost
Lineage visual · proof schematic build campaign

Proof: which workers did the best work, and what it cost

Proof: which workers did the best work, and what it cost

Every model that has taken a seat in Auteur's own construction, ranked from the receipts. This is the dataset behind the governed build-campaign demonstration on the home page, and it is the same ledger the studio reads when it picks a seat for the next round.

Two words carry the whole table. Finished means an attempt ended done and passed the tests its task declared; nothing else counts. Tries counts every attempt, including the ones verification stopped. A caught failure is the product working, so the low rows stay on the page.

Paid seats are metered in dollars by their provider, to the cent. Subscription seats are the flat-rate plans you already carry, so there is no dollar to show for a task and we do not invent one: they are measured in tasks, minutes and new tokens instead.

Receipts, not vibes · measured 2026-09-02

Paid seats · metered in dollars43of 109 tries finished$37.44 reported spend, to the cent, per seat, per round. Cost per finished task is withheld where a try did not report its spend: a partial sum is never divided.
Subscription seats · flat-rate capacity1,152of 1,512 tries finishedClaude, Codex and Antigravity plans you already pay for. Measured in tasks, minutes and new tokens, never dollars, because the meter is the subscription, not the token.
Local seats · your own hardwareProven in the writing pipelinebuild campaigns not yet run locallyLocal models already carry real work in the writing pipeline: Gemma 4 31B drafted a 50,000-word non-fiction guide as first-draft writer, and Gemma 4 12B runs as the background classifier. Unpriced: GPU seconds, not dollars.
WorkerTriesFinishedPass rateMetered spendSpend reportedCost per finished task
Qwen3.8 Maxqwen/qwen3.8-max · high effort262077%$13.8224 / 26not enough data
GLM 5.2z-ai/glm-5.2 · high effort642133%$14.1136 / 64not enough data
Kimi K3moonshotai/kimi-k3 · high effort5120% few tries$4.924 / 5not enough data
Grok 4.6x-ai/grok-4.6 · high effort7114% few tries$4.606 / 7not enough data
Gemini 3.5 Flashgoogle/gemini-3.5-flash · high effort700% few tries0 / 7no price reported

Why "not enough data" on a paid row: a cost per finished task is only shown when every try reported its spend. A silent try makes the ratio a guess, and guesses are not published as measurements. Why the low pass rates are on the page: a caught failure is the product working. Those tries were stopped by verification before anything shipped.

Subscription seats, measured in tasks, minutes and tokens

WorkerTriesFinishedPass rateMinutes per finished taskNew tokens per jobDollars
Claude Opus 4.8claude-opus-4-8 · high effort2121100%9.0470kflat-rate
Claude Sonnet 5claude-sonnet-5 · high effort14212689%11.4430kflat-rate
Claude Sonnet 5claude-sonnet-5 · xhigh effort625487%16.2571kflat-rate
Claude Opus 5claude-opus-5 · xhigh effort13611685%27.4562kflat-rate
Codex, GPT-5.6 Solgpt-5.6-sol · max effort231983%23.2179kflat-rate
Claude Opus 5claude-opus-5 · high effort1199782%20.6411kflat-rate
Gemini 3.7 Flash, via Antigravitygemini-3.7-flash · high effort302480%12.5311kflat-rate
Codex, GPT-5.6 Solgpt-5.6-sol · medium effort151280% few tries7.091kflat-rate
Codex, GPT-5.6 Solgpt-5.6-sol · high effort33625977%14.8133kflat-rate
Claude Haiku 4.5claude-haiku-4-5 · high effort4375% few tries7.2258kflat-rate
Codex, GPT-5.6 Lunagpt-5.6-luna · high effort947074%10.4134kflat-rate
Codex, GPT-5.6 Solgpt-5.6-sol · xhigh effort23416571%22.1203kflat-rate
Earlier or unlisted seatsmodel version not recorded966871%16.4341kflat-rate
Codex, GPT-5.6 Terragpt-5.6-terra · medium effort7571% few tries10.976kflat-rate
Gemini 3.6 Flash, via Antigravitygemini-3.6-flash · high effort916470%6.9flat-rate
Claude Haiku 4.5claude-haiku-4-5 · medium effort6467% few tries11.3266kflat-rate
Gemini 3.7 Flash, via Antigravitygemini-3.7-flash · medium effort9556% few tries17.0flat-rate
Gemini 3.1 Pro, via Antigravitygemini-3.1-pro · high effort552749%16.9522kflat-rate
Claude Opus 5claude-opus-5 · max effort301343%61.5274kflat-rate
Codex, GPT-5.6 Lunagpt-5.6-luna · xhigh effort200% few tries141kflat-rate

"New tokens" is what the model read or wrote for the first time. Cache re-reads are the provider's discount, not our work, so they are not counted. Rows under 20 tries are dimmed: a perfect score from a handful of attempts is a hint, not a fact.

Pass rate by worker, with how many tries it rests on

Bar length is the share of tries that finished and passed their tests. The number after each bar is the try count; only workers with 20 or more tries are charted. Shorter bars are not hidden: they are the proof that verification catches what a model gets wrong.

0%25%50%75%100%Claude Opus 4.8 · high100% · 21Claude Sonnet 5 · high89% · 142Claude Sonnet 5 · xhigh87% · 62Claude Opus 5 · xhigh85% · 136Codex, GPT-5.6 Sol · max83% · 23Claude Opus 5 · high82% · 119Gemini 3.7 Flash, via Antigravity · high80% · 30Codex, GPT-5.6 Sol · high77% · 336Qwen3.8 Max · high77% · 26Codex, GPT-5.6 Luna · high74% · 94Codex, GPT-5.6 Sol · xhigh71% · 234Gemini 3.6 Flash, via Antigravity · high70% · 91Gemini 3.1 Pro, via Antigravity · high49% · 55Claude Opus 5 · max43% · 30GLM 5.2 · high33% · 64
paid seat, metered in dollarssubscription seat, flat-rate, no dollar figure

Read from the operator dashboard on 2026-09-02. Live counters; they keep moving as builds complete.

How to read this honestly

More effort is not more finish. The same model at its highest reasoning setting finished fewer of its tries and took three times as long as it did at a high setting, and the same pattern holds across providers. Auteur's seat selection is drawn from rows like these, not from a leaderboard someone else ran on a benchmark that is not your work.

A cost per finished task is only shown when every try reported its spend. A partial sum divided by a full count would understate the true cost, so the column reads "not enough data" until the coverage is complete.

Cache re-reads are not counted as work anywhere on this page. When a provider serves a token from cache it is that provider's discount, on every tool that uses the same model, so it belongs to them and not to us.

How trust works here · OpenRouter seats · The eval ledger