Charge — describe the failure
The forge
The best benchmarks come from a specific itch. Start from a failure you have actually watched happen, then write the assertions that would catch it — without asking another model whether the answer was acceptable.
A task is a failure you have actually seen. Write the prompt that provokes it, then write assertions that decide the answer without asking a model whether it was good.
Assertions
determinism 100%assertion 1
every field is validated before it is stored
Charge — what the grade will charge you for
The moment you save, the engine runs. These are the six weights it will apply, and they are published in every API response and every export so the number is never a black box.
| Factor | Weight | How to raise it |
|---|---|---|
| determinism | 0.26 | Replace the heaviest judge_rubric with a regex or a JSON path. |
| discrimination | 0.22 | Record a weaker model, or add a near-miss fixture. |
| fixture seal | 0.18 | Pin the revision, the seed, the temperature, and declare each input. |
| assertion specificity | 0.16 | Convert substring assertions into regex or JSON paths. |
| reproduction | 0.10 | Name the model and pin its revision; prefer an ungated one. |
| cost fit | 0.08 | Shorten the prompt, or raise the budget to the measured figure. |
| sum | 1.00 | Published, and asserted in the test suite. |