# Evals — pick agents and models with your own requests

> Re-run one role of an already merged request with another agent or model and have a judge score it against what was actually done. Rubrics, @eval, cost and limits.

URL: https://www.ghosty.studio/en/docs/factory/evals

Public benchmarks don't tell you how a model does **in your repo**. An eval takes a request that was already merged, re-runs **one role** with another agent or model from the same starting point, and a judge scores it against what was actually done. With a few of them you pick each role's model with data.

## Run one

In `/factory` → **Closed** requests → 🧪 on the request → pick **role**, **agent** and/or **model** → "Run eval". Workspace owner only (it spends tokens from their key). The eval opens its own thread in the room so you can see what the agent did; it doesn't count in the request numbers.

| Evaluated role | Starts from | Compared with |
|---|---|---|
| `@plan` | The original request, without seeing the signed plan. | The plan that was signed. |
| `@build` | The same signed plan, on branch `ghosty-eval/<n>` created at the original PR's **base commit**. It opens no PR and the branch is deleted once scored. | The PR that was merged. |
| `@check` | The original PR against the signed plan. It doesn't comment on GitHub. | The first human review and the original `@check`'s. |

## Rubrics

The judge gives a whole number from 1 to 5 per criterion and says whether the result is **worse, same or better** than the original.

| Role | Criteria |
|---|---|
| `@plan` | Covers the request · Concrete · Risks · Quick to sign |
| `@build` | Meets the plan · Tests · Maintainable · Scope |
| `@check` | Finds what matters · No noise · Actionable · Verdict |

The result lands in the thread as a table, with the **model that ran** and its **cost**. The "Evals" table in `/factory` summarizes it by role and configuration: score, worse/same/better, time, role cost and judge cost (medians).

## The judge: `@eval`

By default `@check` judges. A **fixed judge** is better: in "Workspace team" pick an agent for `@eval`.

- Your evals stay comparable over time even if you change `@check`'s model.
- A role is never judged by itself. That's why **evaluating `@check` requires `@eval`**, and only the judge fixed when the eval started can score it.
- A cheaper one from another engine is enough: it scores, it doesn't code.

## Cost

To keep an eval cheap:

- The judge gets everything in its brief (plan, trimmed diffs, rubric) and doesn't explore the repo.
- An identical eval (same request, role, agent and model) isn't repeated: you're told which one exists.
- `@build` and `@check` evals only run on PRs of up to **800 lines**.

For reference, on a small request: a `@plan` eval cost $0.49 and a `@check` one $1.28 for the evaluated role. **The judge can cost as much or more** (with a high-end model, scoring that `@check` cost $1.87 because it reads both full diffs): pick a cheap model for `@eval`. The "Judge" column in the table tells you.
