Evals — pick agents and models with your own requests
Updated 2026-09-28
Read https://www.ghosty.studio/en/docs/factory/evals.md and explain how I compare whether @build does better with another model on my Software Factory requests.
Public benchmarks don't tell you how a model does in your repo. An eval takes a request that was already merged, re-runs one role with another agent or model from the same starting point, and a judge scores it against what was actually done. With a few of them you pick each role's model with data.
Run one
In /factory → Closed requests → 🧪 on the request → pick role, agent and/or model → "Run eval". Workspace owner only (it spends tokens from their key). The eval opens its own thread in the room so you can see what the agent did; it doesn't count in the request numbers.
Rubrics
The judge gives a whole number from 1 to 5 per criterion and says whether the result is worse, same or better than the original.
The result lands in the thread as a table, with the model that ran and its cost. The "Evals" table in /factory summarizes it by role and configuration: score, worse/same/better, time, role cost and judge cost (medians).
The judge: @eval
By default @check judges. A fixed judge is better: in "Workspace team" pick an agent for @eval.
- Your evals stay comparable over time even if you change
@check's model. - A role is never judged by itself. That's why evaluating
@checkrequires@eval, and only the judge fixed when the eval started can score it. - A cheaper one from another engine is enough: it scores, it doesn't code.
Cost
To keep an eval cheap:
- The judge gets everything in its brief (plan, trimmed diffs, rubric) and doesn't explore the repo.
- An identical eval (same request, role, agent and model) isn't repeated: you're told which one exists.
@buildand@checkevals only run on PRs of up to 800 lines.
For reference, on a small request: a @plan eval cost $0.49 and a @check one $1.28 for the evaluated role. The judge can cost as much or more (with a high-end model, scoring that @check cost $1.87 because it reads both full diffs): pick a cheap model for @eval. The "Judge" column in the table tells you.