Software Factory

Evals — pick agents and models with your own requests

Re-run one role of an already merged request with another agent or model and have a judge score it against what was actually done. Rubrics, @eval, cost and limits.

Updated 2026-09-28

Ask your coding agent

Read https://www.ghosty.studio/en/docs/factory/evals.md and explain how I compare whether @build does better with another model on my Software Factory requests.

Paste it in Claude Code, Cursor or Codex with the ghosty-agent skill installed.

Public benchmarks don't tell you how a model does in your repo. An eval takes a request that was already merged, re-runs one role with another agent or model from the same starting point, and a judge scores it against what was actually done. With a few of them you pick each role's model with data.

Run one

In /factory → Closed requests → 🧪 on the request → pick role, agent and/or model → "Run eval". Workspace owner only (it spends tokens from their key). The eval opens its own thread in the room so you can see what the agent did; it doesn't count in the request numbers.

Evaluated roleStarts fromCompared with
@planThe original request, without seeing the signed plan.The plan that was signed.
@buildThe same signed plan, on branch ghosty-eval/<n> created at the original PR's base commit. It opens no PR and the branch is deleted once scored.The PR that was merged.
@checkThe original PR against the signed plan. It doesn't comment on GitHub.The first human review and the original @check's.

Rubrics

The judge gives a whole number from 1 to 5 per criterion and says whether the result is worse, same or better than the original.

RoleCriteria
@planCovers the request · Concrete · Risks · Quick to sign
@buildMeets the plan · Tests · Maintainable · Scope
@checkFinds what matters · No noise · Actionable · Verdict

The result lands in the thread as a table, with the model that ran and its cost. The "Evals" table in /factory summarizes it by role and configuration: score, worse/same/better, time, role cost and judge cost (medians).

The judge: @eval

By default @check judges. A fixed judge is better: in "Workspace team" pick an agent for @eval.

  • Your evals stay comparable over time even if you change @check's model.
  • A role is never judged by itself. That's why evaluating @check requires @eval, and only the judge fixed when the eval started can score it.
  • A cheaper one from another engine is enough: it scores, it doesn't code.

Cost

To keep an eval cheap:

  • The judge gets everything in its brief (plan, trimmed diffs, rubric) and doesn't explore the repo.
  • An identical eval (same request, role, agent and model) isn't repeated: you're told which one exists.
  • @build and @check evals only run on PRs of up to 800 lines.

For reference, on a small request: a @plan eval cost $0.49 and a @check one $1.28 for the evaluated role. The judge can cost as much or more (with a high-end model, scoring that @check cost $1.87 because it reads both full diffs): pick a cheap model for @eval. The "Judge" column in the table tells you.