audit
thinkthen audit RESULTS KEY grades saved answers against answers you already know. It sends no request and reads no API key.
What it reads
RESULTS holds saved result lines, as --details writes them, from any of the ten functions. KEY holds the right answer for some of the records, matched by id. --id POINTER names another field. RESULTS may be - for standard input.
What it prints
One JSON object for each question, one line each. The counts come first:
rows: the answers that did not fail.failed: the answers that did.labeledandunlabeled: the answers the key covers, and the rest.threshold: the rule the answers were read under.rightandwrong: the labeled answers that match the key, and those that do not.unsure: the labeled answers that were not sure under the rule. They count as neither right nor wrong.tied: the labeledchooseorfindanswers where two or more options share the top. Afindanswer counts as tied only whennoneis among them. A tie is tied under every rule. It stays apart fromunsure.tied_holding_keycounts the ties where the right answer was one of the tied options.
Further measures follow, then suggested, a threshold tuned on one part of the records and checked on the other. The audit specification defines every member. --table prints the same results for a person.
thinkthen audit shown.jsonl shown-key.jsonl \
--threshold 0.2:0.8 |
jq '{rows, right, wrong, unsure, tied}'{
"rows": 10,
"right": 3,
"wrong": 2,
"unsure": 5,
"tied": 0
}Any bar, at no cost
--threshold reads the saved probabilities under another cut or band. No request is sent, because the probabilities are already saved. The Beatles audit page grades the same ten songs at three bars. --write FILE writes the suggested threshold, when it holds steady, into the question file the answers came from.
diff · Test it before you trust it · The audit specification
On GitHub: github.com/botassembly/thinkthen