Reference task · read-only
CSV unit-of-measure drift
reference4 assertions3 transcriptsheat 0%
The model returns micrograms where the canonical schema demands milligrams, and the downstream loader silently reads the wrong column.
You are given three rows from a chemistry export.
Return ONLY a JSON object of the form:
{"rows":[{"id":string,"amount":number,"unit":"mg"}],"meta":{"units":"mg","count":number}}
Convert every amount to milligrams. Do not add commentary.| Target model | google/gemini-2.5-flash |
| Revision | seed-reference-revision |
| Seed | 20260923 |
| Temperature | 0 |
| Token budget | 400 |
| Declared inputs | 2 sealed |
- export_rows.csv=sha256:6f1a2c9d4e8b73a5109fd2c4e7b81a3d6c5f0e2b94d7a1c8e3f5b0d2a6c9e4f7
- schema_v3.json=sha256:b3d9e1f04c7a2685e93b1d0c7f4a62e8d5c0b93f7a1e6d4c8b2f0a3e7d5c9b16
google/gemini-2.5-flash — Not published on the Hugging Face Hub, so there is no public revision to pin. A reader cannot reproduce this row from a revision hash.
Tap — the grade, and the dial that changes it
Total assertion weight handed to language-model judges. At 0 every point of the grade is decided by exact comparison and the task is cold. As it rises, more of the number is an opinion and the panel heats up.
This task has no judge rubric, so there is nothing to heat. Add a judge_rubric assertion to see the effect.
The grade is decided by exact assertions, the recorded models do not score alike, and the inputs are pinned.
9 of 9 assertion weight is decided without a model
population sigma 0.3667 over 3 recorded models against a target of 0.22
revision, seed, temperature and every declared input are pinned
weight-averaged specificity 0.933 over 4 assertions
target model named, revision pinned, seed pinned, temperature pinned
mean completion 95.3 tokens is inside the 400 budget
Weakest factor: assertion specificity. Your assertions mostly match substrings. Convert the heaviest ones to a regex or a JSON path.
Tape — every assertion, with the span that decided it
rank 1determinism 1.001840 ms96 tokens out
meta.units = mg
meta.count = 3
absent as required: "µg"
465 lies in [400, 500]
Recorded completion
{"rows":[{"id":"a","amount":120,"unit":"mg"},{"id":"b","amount":250,"unit":"mg"},{"id":"c","amount":95,"unit":"mg"}],"meta":{"units":"mg","count":3,"total":465}}Bundled reference material, not a live model run.
Decide
Bundled reference tasks are read-only so the example stays stable. Decisions you record live on your own tasks.
Ingot — take it with you
A file you can keep: the factor table, every transcript grade, the sealed inputs, the provenance and the head of the chain.
SHA-384 over canonical JSON, one chain per task, chained from a fixed genesis. Recompute it yourself — this endpoint returns the first broken link if there is one.
| Links checked | 0 |
| Replay | intact |
| Head | crucible/v1/genesis… |