ThinkThen

Reference

Recording and the answer cache

ThinkThen saves every good answer, one per question. Ask the same question again and it reads the saved answer instead of sending a request. Caching and replay shows it at work.

--plan

--plan reads the whole input and prints the first request it would send, then one summary line. It sends nothing and needs no key. The Backends page shows the request.

The plan ends with a summary line: one record, one request, and the size of what it would send.
question="Does the customer ask for a refund?"

printf '%s\n' "I want to send this back." |
thinkthen decide "$question" --plan |
tail -1 |
jq .
Output
{
  "records": 1,
  "requests": 1,
  "estimated_bytes": 211,
  "estimated_input_tokens": {
    "lower": 108,
    "upper": 192
  },
  "upper_bound": false
}
exit 0

The old flag --dry-run is refused.

The old flag is refused at exit 2, and the message names the new one.
question="Does the customer ask for a refund?"

printf '%s\n' "I want to send this back." |
thinkthen decide "$question" --dry-run
Output
thinkthen: --dry-run was renamed --plan
exit 2: usage or input error

cache prune --dry-run is a separate flag. It still exists. It lists the answers a prune would remove.

The answer cache

The cache is on by default. The Configuration page says where its folder lives on Linux and macOS, and how THINKTHEN_CACHE moves it.

What makes two questions the same

Each saved answer has a key. The key is the SHA-256 of five things: the adapter, the address, the model, the shared state, and the question as sent. The question as sent includes the quoted evidence. So hello\n and hello\r\n make two keys, and a cosmetic change to the text misses. In a stream, the framing may strip a line end or re-encode parsed JSON before the question is built. Answers from two addresses never mix, because the address is part of every key.

The answer cache keeps its answers in thinkthen.sqlite. A folder committed for replay keeps the same rows in thinkthen.jsonl. A saved answer holds the question and the reply's answer, never a header. No API key can reach it. Only a good answer is saved. A failure is never saved. A later run asks only the questions that failed.

The modes

ModeLooks upSendsSaves
--replay DIRyesnothing; a miss exits 5no
--record DIRnoevery questionyes, replacing a held answer
--cache DIR, THINKTHEN_CACHE, or the default cacheyesthe questions it missesyes
a cache with --refresh-cache, or the model jev-latestnoevery questionyes, replacing a held answer
--no-cachenoevery questionno

--record sends every question, even one the folder already holds, and may charge for it again. Use --cache DIR to reuse what a folder holds. An explicit --record or --replay replaces the default cache for that run. The recording specification holds the full table.

A replay miss

--replay DIR answers from the folder alone, with no key and no network. A question the folder lacks exits 5, and the message names its key.

The folder holds one answer, and this question is not it. The run exits 5 and names the missing key.
question="Does the customer ask for a refund?"

printf '%s\n' "My order came a day early." |
thinkthen decide "$question" \
  --replay recording
Output
thinkthen: the decide request for one document: the replay folder holds no answer for question `ebf765be4444b5dd23230601020a743dbfe11c21c738e5f36307e20b95dbf084`; the key is the SHA-256 of the adapter, address, model, shared state and question as sent
exit 5: a local failure

--facts

--facts prints one thinkthen.run/1 line last on standard error. It counts the records finished, the requests sent and the retries. It always gives seconds, the time the run took. It names the model when every reply named the same one. It adds token counts when every live reply reported them. cache_answers counts answers from the answer cache, never from a replay or record folder.

The answer came from a replay folder. The facts line counts one record, no request sent and no cache answer. Seconds change on every run. jq drops them.
question="Does the customer ask for a refund?"

cat <<'EOF' |
I renewed once this morning, but my card shows two charges.
Please refund the duplicate.
EOF
thinkthen decide "$question" \
  --replay recording \
  --facts 2> facts.json
jq 'del(.seconds)' facts.json
Output
true
{
  "schema": "thinkthen.run/1",
  "records": 1,
  "requests_sent": 0,
  "retries": 0,
  "cache_answers": 0,
  "model": "jev-1.13.0"
}
exit 0

Batching

A stream of records packs many questions into each request. The default batch is max. It fills each request to the backend's limits. --batch 1 sends one record a request. The batch does not change a key. An answer saved under one batch setting replays under another.

Batching may change an answer slightly. Do not batch when you want steadier answers. ThinkThen measured the effect on annotate, where questions share one request. Six borderline questions shared a request with other questions. One answer moved from no to not sure. The largest change in a probability was 0.04. The annotate specification records it.

The same request sent twice can also return another probability. A borderline answer, between 0.33 and 0.67, moved by up to 0.08 in ThinkThen's measurement. --batch 1 does not remove that. A saved answer does. Replay, or a cache with a pinned model, returns the same answer every time. The recording specification records the measurement.

Pruning

  • Nothing trims the cache on its own. It grows, and keeps the judged text, until you run thinkthen cache prune DIR.
  • DIR is always explicit. The prune trims the folder to 100,000,000 bytes, or to --max-size BYTES for one run.
  • --older-than and --answered-by-other-than MODEL select answers to remove first. Either one alone is enough to select an answer.
  • MODEL is the version that answered, as a result's meta.model shows it. Prune refuses the alias passed to --model, or a model no reply in the folder names, and removes nothing. An empty folder accepts any model name.
  • A saved question holds the evidence. ThinkThen creates thinkthen.sqlite with mode 0600, so only its owner can read it.