Skip to content

Lineup — recorded transcripts, re-scored

The lineup

Every row below was produced by crucible-grader-v1.0.0, the same function the API and the agent tools call. Nothing is inherited from a leaderboard elsewhere.

Retry storm under a flaky toolsigma 0.0000mean 1.0000cannot discriminate
#ModelScorePass / fail / openLatencyTokens
1anthropic/claude-sonnet-4
1.00003 / 0 / 01602 ms36
2google/gemini-2.5-flash
1.00003 / 0 / 01310 ms34
3openai/gpt-4.1
1.00003 / 0 / 01500 ms35

Ranked by crucible-grade-v1.0.0. If every row above is identical, the task is measuring agreement rather than capability — the engine charges that as discrimination and the verdict says so.

Open the task

Can anyone reproduce this?

A score you cannot rerun is a rumour. These are live facts from the Hugging Face Hub about the models above: whether the weights are gated, and whether a pinnable revision exists at all.

live · 2026-10-04
ModelReproducible?RevisionLicenseDownloads
Qwen/Qwen2.5-72B-Instructpinnable495f39366efe…other284,503
meta-llama/Llama-3.1-70B-Instructgated1605565b47bb…llama3.199,597
mistralai/Mistral-7B-Instruct-v0.3pinnablec170c708c41d…apache-2.02,147,029
deepseek-ai/DeepSeek-V3pinnablee815299b0bcb…—1,353,839
google/gemini-2.5-flashno public repo
Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
———
anthropic/claude-sonnet-4no public repo
Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
———
openai/gpt-4.1no public repo
Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
———

Model metadata from the Hugging Face Hub public API. Downloads and likes are Hub counters, not usage telemetry. Counter values are Hub totals, not usage telemetry.