Skip to content

Charge

Charge the failure.
Pour the task.

A benchmark that cannot be re-run is a rumour. Crucible grades a task with exact assertions instead of a vibes-based judge, measures whether that task can actually tell two models apart, and files every edit into a hash chain you can replay.

ENGINEcrucible-grade-v1.0.0Six weighted factors summing to 1.00, each produced from a measured quantity with the sentence that produced it attached.
GRADERexact assertions onlyString, regex, JSON-path and numeric-range comparison against the recorded completion. No language model is asked whether an answer is good, so a score can be replayed.

Mould — three reference tasks, three verdicts

The three tasks below ship with the app. Their transcripts are bundled reference material, not live model runs — but the grades are computed live by the real engine every time this page renders.

Retry storm under a flaky toolneeds_hardening

When a tool call fails twice, the model retries in a tight loop instead of surfacing the error. I want to know which models get this right.

69.6
Needs hardeningcrucible-grade-v1.0.0
6 factors
Determinism×0.26 → 0.2600

6 of 6 assertion weight is decided without a model

Discrimination×0.22 → 0.0000

population sigma 0 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.1260

not pinned: revision

Assertion specificity×0.16 → 0.1600

weight-averaged specificity 1 over 3 assertions

Reproduction×0.10 → 0.0700

target model named, revision unpinned, seed pinned, temperature pinned

Cost fit×0.08 → 0.0800

mean completion 35 tokens is inside the 300 budget

Weakest factor: discrimination. Your recorded models score alike. Record a weaker model or add a near-miss fixture so the task can tell them apart.

Inspect

Nested envelope collapsejudge_bound

Given a tool result, the model wraps the payload in a second envelope and the caller's parser sees data: null.

56.7
Judge-boundcrucible-grade-v1.0.0
6 factorsdegraded

57.1% of the assertion weight needs a language-model judge, so part of this score is a claim, not a measurement.

Determinism×0.26 → 0.1114

3 of 7 assertion weight is decided without a model

Discrimination×0.22 → 0.1572

population sigma 0.1572 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.0810

not pinned: revision, temperature

Assertion specificity×0.16 → 0.0869

weight-averaged specificity 0.543 over 3 assertions

Reproduction×0.10 → 0.0500

target model named, revision unpinned, seed pinned, temperature unpinned

Cost fit×0.08 → 0.0800

mean completion 37 tokens is inside the 200 budget

Weakest factor: determinism. Convert the highest-weighted judge rubric into an exact assertion before this task is worth publishing.

Inspect

CSV unit-of-measure driftpublishable

The model returns micrograms where the canonical schema demands milligrams, and the downstream loader silently reads the wrong column.

98.9
Publishablecrucible-grade-v1.0.0
6 factors
Determinism×0.26 → 0.2600

9 of 9 assertion weight is decided without a model

Discrimination×0.22 → 0.2200

population sigma 0.3667 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.1800

revision, seed, temperature and every declared input are pinned

Assertion specificity×0.16 → 0.1493

weight-averaged specificity 0.933 over 4 assertions

Reproduction×0.10 → 0.1000

target model named, revision pinned, seed pinned, temperature pinned

Cost fit×0.08 → 0.0800

mean completion 95.3 tokens is inside the 400 budget

Weakest factor: assertion specificity. Your assertions mostly match substrings. Convert the heaviest ones to a regex or a JSON path.

Inspect

Tap — what a grade actually measures

Published factor weights and what each one measures
FactorWeightWhat it measures
determinism0.26How much of the grade survives without a model in the loop.
discrimination0.22Measured separation between recorded models. A task everything passes measures nothing.
fixture seal0.18Whether the revision, seed, temperature and inputs are pinned.
assertion specificity0.16Regex and JSON paths outrank substring matching.
reproduction0.10Whether a third party can rerun this at all.
cost fit0.08Whether recorded runs fit the declared token budget.

Bands

0–20Unusable
Do not publish. Add deterministic assertions before anything else.
20–40Non-discriminating
Add a near-miss fixture so a plausible wrong answer starts to cost points.
40–60Judge-bound
Convert the highest-weighted judge rubric into an exact assertion.
60–80Needs hardening
Pin the loose input and tighten the lowest-weighted substring assertion.
80–100Publishable
Ship it. Push the task and run it on the Kaggle model proxy.

Ingot — live signals behind the suite

Hugging Face Hub model factslive · 2026-10-04

4 of 4 lineup models expose a pinnable revision. 1 are gated. The rest publish no public repository at all — which means nobody outside the provider can reproduce a pinned run against them, whatever a leaderboard claims.

ModelGatedRevisionDownloads
Qwen/Qwen2.5-72B-Instructopen495f39366e…284,503
meta-llama/Llama-3.1-70B-Instructgated1605565b47…99,597
mistralai/Mistral-7B-Instruct-v0.3openc170c708c4…2,147,029
deepseek-ai/DeepSeek-V3opene815299b0b…1,353,839

Not published on the Hub, so no revision is pinnable: google/gemini-2.5-flash, anthropic/claude-sonnet-4, openai/gpt-4.1.

Model metadata from the Hugging Face Hub public API. Downloads and likes are Hub counters, not usage telemetry.

Cool — take it with you

What a visitor can actually do
  1. Forge. Paste a failure mode, write assertions, record transcripts. Every row is written to the production database.
  2. Inspect. Read the assertion tape: every outcome with the exact span or value that decided it.
  3. Turn the dial. Move model dependence and watch the engine re-grade and the ground hue shift.
  4. Export. Download a Markdown, JSON or CSV dossier with the seal chain attached, or the Kaggle bundle that grades identically.

Start forgingAgent consoleStar on GitHub

What this is not. Crucible does not score models and does not predict capability. A high grade means the task is worth publishing. Where a grade would need a language-model judge, the weight is reported as undecided instead of being quietly counted as a pass.

Crucible is MIT licensed. Engine crucible-grade-v1.0.0. The engine, the grader, the audit chain and the agent tools are all in the repository.