Skip to content

Dossier — take the suite with you

The dossier

Every export carries the six factor weights, the per-transcript grades, the sealed inputs, where the data came from, and the head of the audit chain — so a reader can check the number instead of trusting it.

Download

TaskStateMarkdownJSONCSVKaggle bundle
Retry storm under a flaky tool
3 transcripts · 3 assertions
referenceMarkdownJSONCSVfiles
Nested envelope collapse
3 transcripts · 3 assertions
referenceMarkdownJSONCSVfiles
CSV unit-of-measure drift
3 transcripts · 4 assertions
referenceMarkdownJSONCSVfiles
Markdown

The readable one. Factor table, per-transcript grades, sealed inputs, provenance, chain head.

text/markdown
JSON

The whole task plus verdict, grade and replay result. This is what a script should read.

application/json
CSV

One row per recorded transcript, for a spreadsheet. Drops into a leaderboard notebook as-is.

text/csv

What every export contains

  • The six published weights and each factor’s real contribution, so the bars sum to the score.
  • Every assertion outcome with the exact span or value that decided it.
  • Which transcripts are recorded runs and which are bundled reference material, labelled individually.
  • Where external data came from, when it was retrieved, and whether it was live or a sealed snapshot.
  • The chain head plus the command to recompute it: curl https://crucibleforge.vercel.app/api/tasks/<id>/integrity

Caveat carried in every file: a grade is a statement about a task, not about a model. Where a grade would need a language-model judge, the weight is reported as undecided rather than counted as a pass.