# Test it before you trust it

Every answer comes with the model's probability. You set the bar, and you count the misses on your own records.

## Grade it on your own records

Take records you have already labeled. Put your labels in a key file, one per record. `thinkthen audit` grades ThinkThen's saved answers against the key. `thinkthen audit --help` shows how to save them.

*Jev was asked whether each song is on Abbey Road. At the default bar of 0.5, 5 are right. Jev says yes wrongly 5 times.*

```
thinkthen audit shown.jsonl shown-key.jsonl |
jq '{
  songs: .rows,
  right,
  wrong_yes: .false_yes,
  missed_yes: .false_no
}'
```

*Output*

```
{
  "songs": 10,
  "right": 5,
  "wrong_yes": 5,
  "missed_yes": 0
}
```

*exit 0*

The [audit page](/learn/beatles-bench/audit/) grades the same records at other bars.

## The threshold is yours

The model reports a probability. Your threshold turns it into an answer. One number is a cut. Two numbers are a band, and the middle is not sure.

*Under the band 0.2:0.8, the refund request is yes and the thank-you is no. The send-back line could mean an exchange or money back. It lands in the band as null.*

```
question="Does the customer ask for a refund?"

cat <<'EOF' |
Please refund my order. It arrived broken.
Thanks for the quick help yesterday!
I want to send this back.
EOF
thinkthen decide "$question" \
  --batch 1 \
  --lines \
  --threshold 0.2:0.8 |
jq .
```

*Output*

```
{
  "input": "Please refund my order. It arrived broken.",
  "value": true
}
{
  "input": "Thanks for the quick help yesterday!",
  "value": false
}
{
  "input": "I want to send this back.",
  "value": null
}
```

*exit 0*

A band hands a person the calls a pipeline should not make alone. The backend may fail to answer a question. ThinkThen marks that question failed and counts it. A failure never turns into `null`.

## Four outcomes, four colors

The same four colors run through this site and the terminal.

- **yes** exit 0 prints `true`

- **no** exit 1 prints `false`

- **not sure** exit 3 prints `null`

- **broken** exit 2, 4, 5, 6, 7, 70 prints nothing

## The words for the numbers

| Word | What it is |
| --- | --- |
| probability | A number the model reported, passed through unchanged. For yes or no, the chance the answer is yes. For `choose` and `score`, the share given to each option or level. |
| probabilities | The model's full list, one number per option or level, adding to 1. |
| threshold | Your bar on a probability. An input. |
| band | A low bar and a high bar. The middle goes to a person. |
| position | The number `score` prints: where the evidence lands along your levels. ThinkThen computes it as the average of the level probabilities. |
| strength | The number on a recognized name. ThinkThen computes it. It is not a probability. |

ThinkThen prints no confidence of its own. Set your bar on the probability.

## What leaves the machine

- Each request carries the question, its answer choices and their descriptions, your text, and the model name. When a key is set, the request also carries it in the `Authorization` header. Nothing else of yours goes. With `--field`, only the part you point at goes.
- The tool sends your text as it is. It splits nothing and drops nothing.
- `--plan` prints the plan and sends nothing.
- By default, requests go to TypeSafe. The [backends page](/install/backends/) links TypeSafe's terms and what they say about retention.

[Where the key goes](/install/backends/) · [Saved answers and limits](/reference/) · [What it will not do](/refusals/)

On GitHub: [github.com/botassembly/thinkthen](https://github.com/botassembly/thinkthen)
