Skip to content

Reference task · read-only

Nested envelope collapse

reference3 assertions3 transcriptsheat 57%degraded

Back to suite

Given a tool result, the model wraps the payload in a second envelope and the caller's parser sees data: null.

PROMPT UNDER TEST
A tool returned this payload:
{"status":"ok","data":{"rows":[1,2,3]}}
Reply with ONLY the JSON the caller expects, which is the inner data object verbatim.
PINS
Target modelopenai/gpt-4.1
Revisionnot pinned
Seed4242
Temperaturenot pinned
Token budget200
Declared inputs1 sealed
  • tool_result_07.json=sha256:11ac90b7e2d34f6810c5a7be9d2f40318c6e5a9b7d0c2f4e8a1b3d5f7c9e0a24

openai/gpt-4.1 — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.

Tap — the grade, and the dial that changes it

How much does a judge decide?

Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.

43%

Bundled reference tasks are read-only. Forge your own to move the dial.

56.7
Judge-boundcrucible-grade-v1.0.0
6 factorsdegraded

Too much of the grade depends on a language model in the loop, so two runs of the same suite can disagree.

57.1% of the assertion weight needs a language-model judge, so part of this score is a claim, not a measurement.

Determinism×0.26 → 0.1114

3 of 7 assertion weight is decided without a model

Discrimination×0.22 → 0.1572

population sigma 0.1572 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.0810

not pinned: revision, temperature

Assertion specificity×0.16 → 0.0869

weight-averaged specificity 0.543 over 3 assertions

Reproduction×0.10 → 0.0500

target model named, revision unpinned, seed pinned, temperature unpinned

Cost fit×0.08 → 0.0800

mean completion 37 tokens is inside the 200 budget

Weakest factor: determinism. Convert the highest-weighted judge rubric into an exact assertion before this task is worth publishing.

spread 0.1572 · mean 0.5556

Tape — every assertion, with the span that decided it

rank 1determinism 0.431120 ms22 tokens outhas undecided weight

PASS w1rows sit at the top level json_path_equals
rows = [1,2,3]
FAIL w2no nested data envelope regex
no match for /"data"\s*:/i
OPEN w4answer contains no hedging preamble judge_rubric
judge rubric, not decided on this server: "The completion contains no preamble such as 'Sure', 'Certainly' or 'Here is'. Judge on the completion text alone."

Recorded completion

{"rows":[1,2,3]}

Bundled reference material, not a live model run.

Decide

Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.

Ingot — take it with you

Dossier

A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.

Kaggle bundleSuite dossier

Integrity

SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.

Links checked0
Replayintact
Headcrucible/v1/genesis…
Replay via API