Eval
Nothing here is accepted on vibes. Eval is the studio's judgment layer, scoring, receipting, and remembering, so quality is measured instead of assumed.
Judged before accepted
Drafts, builds, shots, and answers are scored, often by a different model family than the one that made them, before they touch the project. The checker is never the builder.
Receipts, always
Every attempt leaves a receipt: what ran, what it cost, what was judged, what happened. Failures are kept, classified, and useful, you always know how a seat performed.
“Score tonight’s three drafts against my last accepted chapter and tell me which one earns a read.”
How eval looks at a picture
A render request never returns one take on faith. Candidates come back and are scored against a rubric, does the face match the character the project knows, does the composition serve the shot, did the image actually follow the prompt, is anything artifacted, and the winner is accepted with its scores on the record. The rejects keep their scores too, which is how the next request starts smarter.
How eval powers the workshop
Eval is the workshop's compass. Every overnight session ends in an eval pass: what cleared the gate is staged for your morning approval, what fell short goes back with notes attached, and both outcomes land in memory. That is the loop that makes session twelve measurably sharper than session one, scored attempts, remembered verdicts, and a curve that bends upward.
The ledger remembers models, too
Eval does not only judge the work; it keeps book on the workers. Every model you run accrues a track record in the studio's ledger: what it is fast at, what it fumbles, what its output really costs against what it delivers. Failure classes are remembered, and the next build routes around them. That record is what lets the classifier rerank seats mid-build, swap a fumbling model for a better one while the campaign is still running, and it is why trusting an inexpensive open-source model here is a measured decision, not a leap of faith: it has a scorecard.
It learns your taste
Eval outcomes feed memory, and memory shapes the next run. The system gets measurably better at your preferences over time, in code, prose, or picture.