Charge
Charge the failure.
Pour the task.
A benchmark that cannot be re-run is a rumour. Crucible grades a task with exact assertions instead of a vibes-based judge, measures whether that task can actually tell two models apart, and files every edit into a hash chain you can replay.
Mould — three reference tasks, three verdicts
The three tasks below ship with the app. Their transcripts are bundled reference material, not live model runs — but the grades are computed live by the real engine every time this page renders.
When a tool call fails twice, the model retries in a tight loop instead of surfacing the error. I want to know which models get this right.
6 of 6 assertion weight is decided without a model
population sigma 0 over 3 recorded models against a target of 0.22
not pinned: revision
weight-averaged specificity 1 over 3 assertions
target model named, revision unpinned, seed pinned, temperature pinned
mean completion 35 tokens is inside the 300 budget
Weakest factor: discrimination. Your recorded models score alike. Record a weaker model or add a near-miss fixture so the task can tell them apart.
Given a tool result, the model wraps the payload in a second envelope and the caller's parser sees data: null.
57.1% of the assertion weight needs a language-model judge, so part of this score is a claim, not a measurement.
3 of 7 assertion weight is decided without a model
population sigma 0.1572 over 3 recorded models against a target of 0.22
not pinned: revision, temperature
weight-averaged specificity 0.543 over 3 assertions
target model named, revision unpinned, seed pinned, temperature unpinned
mean completion 37 tokens is inside the 200 budget
Weakest factor: determinism. Convert the highest-weighted judge rubric into an exact assertion before this task is worth publishing.
The model returns micrograms where the canonical schema demands milligrams, and the downstream loader silently reads the wrong column.
9 of 9 assertion weight is decided without a model
population sigma 0.3667 over 3 recorded models against a target of 0.22
revision, seed, temperature and every declared input are pinned
weight-averaged specificity 0.933 over 4 assertions
target model named, revision pinned, seed pinned, temperature pinned
mean completion 95.3 tokens is inside the 400 budget
Weakest factor: assertion specificity. Your assertions mostly match substrings. Convert the heaviest ones to a regex or a JSON path.
Tap — what a grade actually measures
| Factor | Weight | What it measures |
|---|---|---|
| determinism | 0.26 | How much of the grade survives without a model in the loop. |
| discrimination | 0.22 | Measured separation between recorded models. A task everything passes measures nothing. |
| fixture seal | 0.18 | Whether the revision, seed, temperature and inputs are pinned. |
| assertion specificity | 0.16 | Regex and JSON paths outrank substring matching. |
| reproduction | 0.10 | Whether a third party can rerun this at all. |
| cost fit | 0.08 | Whether recorded runs fit the declared token budget. |
Bands
| 0–20 | Unusable Do not publish. Add deterministic assertions before anything else. |
| 20–40 | Non-discriminating Add a near-miss fixture so a plausible wrong answer starts to cost points. |
| 40–60 | Judge-bound Convert the highest-weighted judge rubric into an exact assertion. |
| 60–80 | Needs hardening Pin the loose input and tighten the lowest-weighted substring assertion. |
| 80–100 | Publishable Ship it. Push the task and run it on the Kaggle model proxy. |
Ingot — live signals behind the suite
4 of 4 lineup models expose a pinnable revision. 1 are gated. The rest publish no public repository at all — which means nobody outside the provider can reproduce a pinned run against them, whatever a leaderboard claims.
| Model | Gated | Revision | Downloads |
|---|---|---|---|
| Qwen/Qwen2.5-72B-Instruct | open | 495f39366e… | 284,503 |
| meta-llama/Llama-3.1-70B-Instruct | gated | 1605565b47… | 99,597 |
| mistralai/Mistral-7B-Instruct-v0.3 | open | c170c708c4… | 2,147,029 |
| deepseek-ai/DeepSeek-V3 | open | e815299b0b… | 1,353,839 |
Not published on the Hub, so no revision is pinnable: google/gemini-2.5-flash, anthropic/claude-sonnet-4, openai/gpt-4.1.
Model metadata from the Hugging Face Hub public API. Downloads and likes are Hub counters, not usage telemetry.
Cool — take it with you
- Forge. Paste a failure mode, write assertions, record transcripts. Every row is written to the production database.
- Inspect. Read the assertion tape: every outcome with the exact span or value that decided it.
- Turn the dial. Move model dependence and watch the engine re-grade and the ground hue shift.
- Export. Download a Markdown, JSON or CSV dossier with the seal chain attached, or the Kaggle bundle that grades identically.
What this is not. Crucible does not score models and does not predict capability. A high grade means the task is worth publishing. Where a grade would need a language-model judge, the weight is reported as undecided instead of being quietly counted as a pass.
Crucible is MIT licensed. Engine crucible-grade-v1.0.0. The engine, the grader, the audit chain and the agent tools are all in the repository.