audit grades a run.
You already know the right answer for some of your records. audit grades saved answers against those answers at any bar. It sends no request.
Here Jev was asked whether each of 10 songs is on Abbey Road, from the title alone. The slide's no-context column shows these answers. The slide reads each answer at the band 0.2:0.8. A red cross marks a wrong answer, and an amber ? marks a not-sure answer. The first example below runs audit on the same 10 songs. The diff page asks about them again with context.
Run it
thinkthen audit shown.jsonl shown-key.jsonl \
--threshold 0.2:0.8 |
jq '{
songs: .rows,
right,
wrong,
not_sure: .unsure
}'{
"songs": 10,
"right": 3,
"wrong": 2,
"not_sure": 5
}Try the default bar
thinkthen audit shown.jsonl shown-key.jsonl |
jq '{
songs: .rows,
right,
wrong_yes: .false_yes,
missed_yes: .false_no
}'{
"songs": 10,
"right": 5,
"wrong_yes": 5,
"missed_yes": 0
}Change the bar
thinkthen audit shown.jsonl shown-key.jsonl \
--threshold 0.78 |
jq '{
songs: .rows,
right,
wrong_yes: .false_yes,
missed_yes: .false_no
}'{
"songs": 10,
"right": 8,
"wrong_yes": 2,
"missed_yes": 0
}The lesson
From the title alone, half the answers fall inside the band. At 0.5, Jev says yes to 5 songs from other albums. At 0.78, only A Day in the Life and The Long and Winding Road remain wrong, and no Abbey Road song is lost.
The answers you already know grade any bar, at no cost.
On GitHub: github.com/botassembly/beatles-bench/tree/main/examples/audit