Skip to content

Reference task · read-only

CSV unit-of-measure drift

reference4 assertions3 transcriptsheat 0%

Back to suite

The model returns micrograms where the canonical schema demands milligrams, and the downstream loader silently reads the wrong column.

PROMPT UNDER TEST
You are given three rows from a chemistry export.
Return ONLY a JSON object of the form:
{"rows":[{"id":string,"amount":number,"unit":"mg"}],"meta":{"units":"mg","count":number}}
Convert every amount to milligrams. Do not add commentary.
PINS
Target modelgoogle/gemini-2.5-flash
Revisionseed-reference-revision
Seed20260923
Temperature0
Token budget400
Declared inputs2 sealed
  • export_rows.csv=sha256:6f1a2c9d4e8b73a5109fd2c4e7b81a3d6c5f0e2b94d7a1c8e3f5b0d2a6c9e4f7
  • schema_v3.json=sha256:b3d9e1f04c7a2685e93b1d0c7f4a62e8d5c0b93f7a1e6d4c8b2f0a3e7d5c9b16

google/gemini-2.5-flash — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.

Tap — the grade, and the dial that changes it

How much does a judge decide?

Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.

100%

This task has no judge rubric, so there is nothing to heat. Add a judge_rubric assertion to see the effect.

98.9
Publishablecrucible-grade-v1.0.0
6 factors

The grade is decided by exact assertions, the recorded models do not score alike, and the inputs are pinned.

Determinism×0.26 → 0.2600

9 of 9 assertion weight is decided without a model

Discrimination×0.22 → 0.2200

population sigma 0.3667 over 3 recorded models against a target of 0.22

Fixture seal×0.18 → 0.1800

revision, seed, temperature and every declared input are pinned

Assertion specificity×0.16 → 0.1493

weight-averaged specificity 0.933 over 4 assertions

Reproduction×0.10 → 0.1000

target model named, revision pinned, seed pinned, temperature pinned

Cost fit×0.08 → 0.0800

mean completion 95.3 tokens is inside the 400 budget

Weakest factor: assertion specificity. Your assertions mostly match substrings. Convert the heaviest ones to a regex or a JSON path.

spread 0.3667 · mean 0.4815

Tape — every assertion, with the span that decided it

rank 1determinism 1.001840 ms96 tokens out

PASS w3every unit is mg json_path_equals
meta.units = mg
PASS w2exactly three rows counted json_path_equals
meta.count = 3
PASS w2no micrograms anywhere not_contains
absent as required: "µg"
PASS w2total mass lands in 400-500 mg number_between
465 lies in [400, 500]

Recorded completion

{"rows":[{"id":"a","amount":120,"unit":"mg"},{"id":"b","amount":250,"unit":"mg"},{"id":"c","amount":95,"unit":"mg"}],"meta":{"units":"mg","count":3,"total":465}}

Bundled reference material, not a live model run.

Decide

Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.

Ingot — take it with you

Dossier

A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.

Kaggle bundleSuite dossier

Integrity

SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.

Links checked0
Replayintact
Headcrucible/v1/genesis…
Replay via API