Introducing ThinkThen
Code sees strings. It can’t tell what they mean.
ThinkThen fixes that. Your code asks one bounded question about some text. It gets back a typed answer and acts on it.
Ask a question, get a value
Here’s a customer message. Does it ask for a refund?
question="Does the customer ask for a refund?"
cat <<'EOF' |
I renewed once this morning, but my card shows two charges.
Please refund the duplicate.
EOF
thinkthen decide "$question"true
The answer is true. Your code reads it as a value. There’s no prose to parse.
Every answer has a shape you choose up front: yes or no, one option from your list, or a number on your scale.
Ten functions you already know
I built ThinkThen around shapes every programmer knows. decide is an if that understands. filter is a grep that understands. rank is a sort that understands.
The other seven fill out the set. choose picks one option. tag names every label that fits. score places text on a scale. find picks the best line. annotate fills out a form. recognize finds names. relate finds how they connect.
The functions page shows each one.
Answers are exit codes
In a script, the answer is also the exit code. Yes exits 0. No exits 1.
Add a band, and the middle gets its own answer. Here the band runs from 0.2 to 0.8. This send-back line lands inside it:
question="Does the customer ask for a refund?"
printf '%s\n' "I want to send this back." |
thinkthen decide "$question" \
--quiet \
--threshold 0.2:0.8
refund_code=$?
case $refund_code in
0) echo "refund it" ;;
1) echo "answer it as usual" ;;
3) echo "send it to a person" ;;
esac
test "$refund_code" -eq 3send it to a person
Not sure exits 3, so a person reads it. A failure gets its own exit code, such as 4 when the backend fails. A failure never looks like an answer.
No training data to start
You don’t train a model. You don’t label a dataset first. The question is plain words.
You still measure before you trust it. Label some real cases, then run thinkthen audit. It grades saved answers against your labels, and you pick the threshold from that. Test it before you trust it shows how.
It runs in your language
ThinkThen is one command-line tool and 24 bindings. Every binding calls the same Rust engine.
The bindings include DuckDB, SQLite and PostgreSQL. In those, a question sits in a SQL WHERE clause like any other condition.
Cheap and fast
On Beatles Bench, 1,000 Jev answers cost about $0.016. That’s from the bench’s cost table.
A live call through a library binding took a median of 138 ms on 2026-10-01. ThinkThen’s own work was about 1 to 2 ms of it. From the command line, each run starts a new process and opens a new connection. A live command run took a median of 219 ms. The overhead page has every measurement.
An open standard
ThinkThen speaks System One, the request format Jev answers. Any System One server works.
Jev, from TypeSafe, is the default. Liquid AI’s d1 works through the Liquid backend. A local model works through Ollama, with no key. thinkthen check tells you whether a server works with ThinkThen.
One caution: a threshold belongs to one model. Switch models, and you measure again.
Test it like ordinary code
ThinkThen can record the answers a run gets. A test replays them later with no network and no key. Caching and replay shows how.
That’s how this site works. Every example on it is a replayed test, and the build fails when one changes.
Where it breaks
Planted facts move answers. In one probe, a false claim added to a customer message moved the probability of yes by 0.57. Jev reads a planted claim as evidence. Use a band, and send the middle to a person.
tag misses part of the set. On Beatles Bench it names exactly the right set of lead singers only 56% of the time. That’s from the bench’s function table. Use tag to fill a queue a person reads.
It only answers. ThinkThen writes no text and takes no action. Your code acts.
Proven in public
We built Beatles Bench to test this in the open. It asks 1,501 questions, most of them about Beatles songs. A script set every right answer from Wikipedia and Wikidata.
Jev gets 70.5% of them right from memory. The model report compares it with other models and with plain search. I build ThinkThen, so weigh that result with care.
Open source
ThinkThen is MIT licensed. I build it at GenomOncology, and the code is on GitHub.
One line installs it:
curl -fsSL https://thinkthen.dev/install.sh | shThe tutorial asks your first question.