---
title: "Proof: which workers did the best work, and what it cost"
canonical_url: https://auteurintelligence.com/proof/build-campaign/
description: "Every model that has taken a seat in Auteur's own construction, ranked from the receipts. This is the dataset behind the governed build-campaign…"
---

# Proof: which workers did the best work, and what it cost

Every model that has taken a seat in Auteur's own construction, ranked from
the receipts. This is the dataset behind the governed build-campaign
demonstration on the [home page](https://auteurintelligence.com/#proof), and it is the same ledger the
studio reads when it picks a seat for the next round.

Two words carry the whole table. **Finished** means an attempt ended done and
passed the tests its task declared; nothing else counts. **Tries** counts
every attempt, including the ones verification stopped. A caught failure is
the product working, so the low rows stay on the page.

Paid seats are metered in dollars by their provider, to the cent. Subscription
seats are the flat-rate plans you already carry, so there is no dollar to show
for a task and we do not invent one: they are measured in tasks, minutes and
new tokens instead.

Receipts, not vibes · measured 2026-09-02

Paid seats · metered in dollars43of 109 tries finished$37.44 reported spend, to the cent, per seat, per round. Cost per finished task is withheld where a try did not report its spend: a partial sum is never divided.

Subscription seats · flat-rate capacity1,152of 1,512 tries finishedClaude, Codex and Antigravity plans you already pay for. Measured in tasks, minutes and new tokens, never dollars, because the meter is the subscription, not the token.

Local seats · your own hardwareProven in the writing pipelinebuild campaigns not yet run locallyLocal models already carry real work in the writing pipeline: Gemma 4 31B drafted a 50,000-word non-fiction guide as first-draft writer, and Gemma 4 12B runs as the background classifier. Unpriced: GPU seconds, not dollars.

Paid seats, metered in dollars

WorkerTriesFinishedPass rateMetered spendSpend reportedCost per finished taskQwen3.8 Maxqwen/qwen3.8-max · high effort262077%$13.8224 / 26not enough dataGLM 5.2z-ai/glm-5.2 · high effort642133%$14.1136 / 64not enough dataKimi K3moonshotai/kimi-k3 · high effort5120% few tries$4.924 / 5not enough dataGrok 4.6x-ai/grok-4.6 · high effort7114% few tries$4.606 / 7not enough dataGemini 3.5 Flashgoogle/gemini-3.5-flash · high effort700% few tries–0 / 7no price reported

Why "not enough data" on a paid row: a cost per finished task is only shown when every try reported its spend. A silent try makes the ratio a guess, and guesses are not published as measurements. Why the low pass rates are on the page: a caught failure is the product working. Those tries were stopped by verification before anything shipped.

Subscription seats, measured in tasks, minutes and tokens

WorkerTriesFinishedPass rateMinutes per finished taskNew tokens per jobDollarsClaude Opus 4.8claude-opus-4-8 · high effort2121100%9.0470kflat-rateClaude Sonnet 5claude-sonnet-5 · high effort14212689%11.4430kflat-rateClaude Sonnet 5claude-sonnet-5 · xhigh effort625487%16.2571kflat-rateClaude Opus 5claude-opus-5 · xhigh effort13611685%27.4562kflat-rateCodex, GPT-5.6 Solgpt-5.6-sol · max effort231983%23.2179kflat-rateClaude Opus 5claude-opus-5 · high effort1199782%20.6411kflat-rateGemini 3.7 Flash, via Antigravitygemini-3.7-flash · high effort302480%12.5311kflat-rateCodex, GPT-5.6 Solgpt-5.6-sol · medium effort151280% few tries7.091kflat-rateCodex, GPT-5.6 Solgpt-5.6-sol · high effort33625977%14.8133kflat-rateClaude Haiku 4.5claude-haiku-4-5 · high effort4375% few tries7.2258kflat-rateCodex, GPT-5.6 Lunagpt-5.6-luna · high effort947074%10.4134kflat-rateCodex, GPT-5.6 Solgpt-5.6-sol · xhigh effort23416571%22.1203kflat-rateEarlier or unlisted seatsmodel version not recorded966871%16.4341kflat-rateCodex, GPT-5.6 Terragpt-5.6-terra · medium effort7571% few tries10.976kflat-rateGemini 3.6 Flash, via Antigravitygemini-3.6-flash · high effort916470%6.9–flat-rateClaude Haiku 4.5claude-haiku-4-5 · medium effort6467% few tries11.3266kflat-rateGemini 3.7 Flash, via Antigravitygemini-3.7-flash · medium effort9556% few tries17.0–flat-rateGemini 3.1 Pro, via Antigravitygemini-3.1-pro · high effort552749%16.9522kflat-rateClaude Opus 5claude-opus-5 · max effort301343%61.5274kflat-rateCodex, GPT-5.6 Lunagpt-5.6-luna · xhigh effort200% few tries–141kflat-rate

"New tokens" is what the model read or wrote for the first time. Cache re-reads are the provider's discount, not our work, so they are not counted. Rows under 20 tries are dimmed: a perfect score from a handful of attempts is a hint, not a fact.

Pass rate by worker, with how many tries it rests on

Bar length is the share of tries that finished and passed their tests. The number after each bar is the try count; only workers with 20 or more tries are charted. Shorter bars are not hidden: they are the proof that verification catches what a model gets wrong.

0%25%50%75%100%Claude Opus 4.8 · high100% · 21Claude Sonnet 5 · high89% · 142Claude Sonnet 5 · xhigh87% · 62Claude Opus 5 · xhigh85% · 136Codex, GPT-5.6 Sol · max83% · 23Claude Opus 5 · high82% · 119Gemini 3.7 Flash, via Antigravity · high80% · 30Codex, GPT-5.6 Sol · high77% · 336Qwen3.8 Max · high77% · 26Codex, GPT-5.6 Luna · high74% · 94Codex, GPT-5.6 Sol · xhigh71% · 234Gemini 3.6 Flash, via Antigravity · high70% · 91Gemini 3.1 Pro, via Antigravity · high49% · 55Claude Opus 5 · max43% · 30GLM 5.2 · high33% · 64paid seat, metered in dollarssubscription seat, flat-rate, no dollar figure

Read from the operator dashboard on 2026-09-02. Live counters; they keep moving as builds complete.

## How to read this honestly

More effort is not more finish. The same model at its highest reasoning
setting finished fewer of its tries and took three times as long as it did at
a high setting, and the same pattern holds across providers. Auteur's seat
selection is drawn from rows like these, not from a leaderboard someone else
ran on a benchmark that is not your work.

A cost per finished task is only shown when every try reported its spend. A
partial sum divided by a full count would understate the true cost, so the
column reads "not enough data" until the coverage is complete.

Cache re-reads are not counted as work anywhere on this page. When a provider
serves a token from cache it is that provider's discount, on every tool that
uses the same model, so it belongs to them and not to us.

[How trust works here](https://auteurintelligence.com/trust/) · [OpenRouter seats](https://auteurintelligence.com/providers/openrouter/) · [The eval ledger](https://auteurintelligence.com/platform/eval/)
