Reference task · read-only
Nested envelope collapse
reference3 assertions3 transcriptsheat 57%degraded
Given a tool result, the model wraps the payload in a second envelope and the caller's parser sees data: null.
A tool returned this payload:
{"status":"ok","data":{"rows":[1,2,3]}}
Reply with ONLY the JSON the caller expects, which is the inner data object verbatim.| Target model | openai/gpt-4.1 |
| Revision | not pinned |
| Seed | 4242 |
| Temperature | not pinned |
| Token budget | 200 |
| Declared inputs | 1 sealed |
- tool_result_07.json=sha256:11ac90b7e2d34f6810c5a7be9d2f40318c6e5a9b7d0c2f4e8a1b3d5f7c9e0a24
openai/gpt-4.1 — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
Tap — the grade, and the dial that changes it
Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.
Bundled reference tasks are read-only. Forge your own to move the dial.
Too much of the grade depends on a language model in the loop, so two runs of the same suite can disagree.
57.1% of the assertion weight needs a language-model judge, so part of this score is a claim, not a measurement.
3 of 7 assertion weight is decided without a model
population sigma 0.1572 over 3 recorded models against a target of 0.22
not pinned: revision, temperature
weight-averaged specificity 0.543 over 3 assertions
target model named, revision unpinned, seed pinned, temperature unpinned
mean completion 37 tokens is inside the 200 budget
Weakest factor: determinism. Convert the highest-weighted judge rubric into an exact assertion before this task is worth publishing.
Tape — every assertion, with the span that decided it
rank 1determinism 0.431120 ms22 tokens outhas undecided weight
rows = [1,2,3]
no match for /"data"\s*:/i
judge rubric, not decided on this server: "The completion contains no preamble such as 'Sure', 'Certainly' or 'Here is'. Judge on the completion text alone."
Recorded completion
{"rows":[1,2,3]}Bundled reference material, not a live model run.
Decide
Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.
Ingot — take it with you
A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.
SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.
| Links checked | 0 |
| Replay | intact |
| Head | crucible/v1/genesis… |