OpenRouter
One key, essentially every public model, and the guardrails to trust the small ones. OpenRouter is a pay-as-you-go gateway to the open model world: frontier names and open-source challengers, big and small, behind a single account with per-token pricing. On its own, that catalog is overwhelming and the cheap models are a gamble. Seated inside Auteur Intelligence™, it becomes something else entirely.
Why you can trust the smaller models here
Our software-engineered guardrails, evaluation, cross-provider review, receipts, and verification gates, keep smaller and open-source models accurate, honest, and fact-checked while they work. That changes the economics of a complex build: work that would once have demanded frontier API spend at every step can run on nimble, inexpensive models for the drafting and the grinding, with the mistakes caught by the scaffolding instead of by your budget. In our own receipted builds, this site included, the metered spend ran to a fraction of what a frontier-only pipeline would have burned.
Save the big guns for where they earn their rate: final-round audit checks at the end of each phase of a coding build, a writing pipeline, or a script-to-image breakdown. Draft local and small, finish frontier.
Every penny tracked, every model benchmarked
Pay-as-you-go means the meter is the truth, and our receipts track it to the penny, per seat, per round, per build. The studio also benchmarks the OpenRouter models you choose against your real work as they run, so your classifier and the engineering around it can reroute a build in progress: a model that starts fumbling gets swapped mid-flight, and the money stops going where the quality is not. Fully cloud-based, no subscription to carry, your key and your terms.
The paid seats, from the receipts · measured 2026-09-02
| Worker | Tries | Finished | Pass rate | Metered spend | Spend reported | Cost per finished task |
|---|---|---|---|---|---|---|
| Qwen3.8 Maxqwen/qwen3.8-max · high effort | 26 | 20 | 77% | $13.82 | 24 / 26 | not enough data |
| GLM 5.2z-ai/glm-5.2 · high effort | 64 | 21 | 33% | $14.11 | 36 / 64 | not enough data |
| Kimi K3moonshotai/kimi-k3 · high effort | 5 | 1 | 20% few tries | $4.92 | 4 / 5 | not enough data |
| Grok 4.6x-ai/grok-4.6 · high effort | 7 | 1 | 14% few tries | $4.60 | 6 / 7 | not enough data |
| Gemini 3.5 Flashgoogle/gemini-3.5-flash · high effort | 7 | 0 | 0% few tries | – | 0 / 7 | no price reported |
Dollars are the provider's own metered figure. A cost per finished task is shown only when every try reported its spend. Subscription seats are not on this table because they are not measured in dollars; the full comparison is the build-campaign proof.
"Run tonight's chapter drafts on the cheapest OpenRouter models that pass our benchmark, route anything that fails review to a frontier seat for the audit, and show me the per-chapter spend in the morning."