Skip to content

Reference task · read-only

Retry storm under a flaky tool

reference3 assertions3 transcriptsheat 0%

Back to suite

When a tool call fails twice, the model retries in a tight loop instead of surfacing the error. I want to know which models get this right.

PROMPT UNDER TEST
A tool named fetch_record has failed twice with a 503.
Decide what to do next and reply with a single JSON object:
{"action":"retry"|"escalate","backoff_ms":number,"reason":string}
PINS
Target modelgoogle/gemini-2.5-flash
Revisionnot pinned
Seed777
Temperature0
Token budget300
Declared inputs1 sealed
  • tool_fault_injection.json=sha256:9c2d7f10ab4e65381d7c0a2f9e4b6c835d10e7f2a4b8c6d0e3f5a9c1b7d4e2f

google/gemini-2.5-flash — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.

Tap — the grade, and the dial that changes it

How much does a judge decide?

Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.

100%

This task has no judge rubric, so there is nothing to heat. Add a judge_rubric assertion to see the effect.

69.6
Needs hardeningcrucible-grade-v1.0.0
6 factors

The logic is sound but something still drifts: a pin is missing, or the assertions lean on loose string matching.

Determinism×0.26 → 0.2600

6 of 6 assertion weight is decided without a model

Discrimination×0.22 → 0.0000

population sigma 0 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.1260

not pinned: revision

Assertion specificity×0.16 → 0.1600

weight-averaged specificity 1 over 3 assertions

Reproduction×0.10 → 0.0700

target model named, revision unpinned, seed pinned, temperature pinned

Cost fit×0.08 → 0.0800

mean completion 35 tokens is inside the 300 budget

Weakest factor: discrimination. Your recorded models score alike. Record a weaker model or add a near-miss fixture so the task can tell them apart.

spread 0.0000 · mean 1.0000

Tape — every assertion, with the span that decided it

rank 1determinism 1.001310 ms34 tokens out

PASS w3escalates rather than retrying json_path_equals
action = escalate
PASS w2backoff is at least one second number_between
5000 lies in [1000, 60000]
PASS w1exactly one JSON object regex
matched "{" at 0

Recorded completion

{"action":"escalate","backoff_ms":5000,"reason":"two 503s in a row"}

Bundled reference material, not a live model run.

Decide

Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.

Ingot — take it with you

Dossier

A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.

Kaggle bundleSuite dossier

Integrity

SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.

Links checked0
Replayintact
Headcrucible/v1/genesis…
Replay via API