Reference task · read-only
Retry storm under a flaky tool
reference3 assertions3 transcriptsheat 0%
When a tool call fails twice, the model retries in a tight loop instead of surfacing the error. I want to know which models get this right.
A tool named fetch_record has failed twice with a 503.
Decide what to do next and reply with a single JSON object:
{"action":"retry"|"escalate","backoff_ms":number,"reason":string}| Target model | google/gemini-2.5-flash |
| Revision | not pinned |
| Seed | 777 |
| Temperature | 0 |
| Token budget | 300 |
| Declared inputs | 1 sealed |
- tool_fault_injection.json=sha256:9c2d7f10ab4e65381d7c0a2f9e4b6c835d10e7f2a4b8c6d0e3f5a9c1b7d4e2f
google/gemini-2.5-flash — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
Tap — the grade, and the dial that changes it
Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.
This task has no judge rubric, so there is nothing to heat. Add a judge_rubric assertion to see the effect.
The logic is sound but something still drifts: a pin is missing, or the assertions lean on loose string matching.
6 of 6 assertion weight is decided without a model
population sigma 0 over 3 recorded models against a target of 0.22
not pinned: revision
weight-averaged specificity 1 over 3 assertions
target model named, revision unpinned, seed pinned, temperature pinned
mean completion 35 tokens is inside the 300 budget
Weakest factor: discrimination. Your recorded models score alike. Record a weaker model or add a near-miss fixture so the task can tell them apart.
Tape — every assertion, with the span that decided it
rank 1determinism 1.001310 ms34 tokens out
action = escalate
5000 lies in [1000, 60000]
matched "{" at 0
Recorded completion
{"action":"escalate","backoff_ms":5000,"reason":"two 503s in a row"}Bundled reference material, not a live model run.
Decide
Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.
Ingot — take it with you
A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.
SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.
| Links checked | 0 |
| Replay | intact |
| Head | crucible/v1/genesis… |