ThinkThen

Reference

audit

thinkthen audit RESULTS KEY grades saved answers against answers you already know. It sends no request and reads no API key.

What it reads

RESULTS holds saved result lines, as --details writes them, from any of the ten functions. KEY holds the right answer for some of the records, matched by id. --id POINTER names another field. RESULTS may be - for standard input.

What it prints

One JSON object for each question, one line each. The counts come first:

  • rows: the answers that did not fail. failed: the answers that did.
  • labeled and unlabeled: the answers the key covers, and the rest.
  • threshold: the rule the answers were read under.
  • right and wrong: the labeled answers that match the key, and those that do not.
  • unsure: the labeled answers that were not sure under the rule. They count as neither right nor wrong.
  • tied: the labeled choose or find answers where two or more options share the top. A find answer counts as tied only when none is among them. A tie is tied under every rule. It stays apart from unsure. tied_holding_key counts the ties where the right answer was one of the tied options.

Further measures follow, then suggested, a threshold tuned on one part of the records and checked on the other. The audit specification defines every member. --table prints the same results for a person.

Ten songs at the band 0.2:0.8: 3 right, 2 wrong, 5 not sure, and no ties.
thinkthen audit shown.jsonl shown-key.jsonl \
  --threshold 0.2:0.8 |
jq '{rows, right, wrong, unsure, tied}'
Output
{
  "rows": 10,
  "right": 3,
  "wrong": 2,
  "unsure": 5,
  "tied": 0
}
exit 0

Any bar, at no cost

--threshold reads the saved probabilities under another cut or band. No request is sent, because the probabilities are already saved. The Beatles audit page grades the same ten songs at three bars. --write FILE writes the suggested threshold, when it holds steady, into the question file the answers came from.