Skip to content

Charge — describe the failure

The forge

The best benchmarks come from a specific itch. Start from a failure you have actually watched happen, then write the assertions that would catch it — without asking another model whether the answer was acceptable.

A task is a failure you have actually seen. Write the prompt that provokes it, then write assertions that decide the answer without asking a model whether it was good.

Assertions

determinism 100%
assertion 1
every field is validated before it is stored

Charge — what the grade will charge you for

The moment you save, the engine runs. These are the six weights it will apply, and they are published in every API response and every export so the number is never a black box.

FactorWeightHow to raise it
determinism0.26Replace the heaviest judge_rubric with a regex or a JSON path.
discrimination0.22Record a weaker model, or add a near-miss fixture.
fixture seal0.18Pin the revision, the seed, the temperature, and declare each input.
assertion specificity0.16Convert substring assertions into regex or JSON paths.
reproduction0.10Name the model and pin its revision; prefer an ungated one.
cost fit0.08Shorten the prompt, or raise the budget to the measured figure.
sum1.00Published, and asserted in the test suite.