Code that knows what you mean
Code sees strings, not meaning. grep finds the word “refund”. It can’t tell whether “I want to send this back” asks for money or for an exchange. Search matches letters. It can’t answer a question about the text.
TypeSafe makes a model called Jev that can. Jev answers a bounded question about the text you hand it. The answer is yes or no, one option from your list, or a place on a scale. Each answer comes with a probability. Jev writes no sentences, so your code has nothing to parse. It needs no training data and no labels. The question goes in as plain words. On Beatles Bench, a thousand answers cost about 0.016 dollars, by the bench’s cost table.
We built ThinkThen around that interface. It’s one program. You pipe in the text and pass one question. ThinkThen prints a bare answer and sets an exit code.
Three answers, three exit codes
Here are three customer messages. The first asks for money back. The second doesn’t. The third could mean either. Ask the same question of each line, with a band from 0.2 to 0.8:
question="Does the customer ask for a refund?"
cat <<'EOF' |
Please refund my order. It arrived broken.
Thanks for the quick help yesterday!
I want to send this back.
EOF
thinkthen decide "$question" \
--batch 1 \
--lines \
--threshold 0.2:0.8 |
jq .{
"input": "Please refund my order. It arrived broken.",
"value": true
}
{
"input": "Thanks for the quick help yesterday!",
"value": false
}
{
"input": "I want to send this back.",
"value": null
}The refund request clears the high bar and answers yes. The thank-you note falls under the low bar and answers no. The send-back line lands inside the band and answers not sure. grep would have said no to that line, and a person would never see it.
A question and its band can live in a file, refund.json:
{
"decide": "Does the customer ask for a refund?",
"true": "The customer asks for money back.",
"false": "Anything else, such as a cancellation or thanks.",
"threshold": "0.2:0.8"
}Ask it of the send-back line:
printf '%s\n' "I want to send this back." |
thinkthen decide @refund.jsonnull
Yes exits 0. No exits 1. Not sure exits 3 and prints null. A script branches on the exit code and parses nothing. A person reviews the middle. The question file diffs in a pull request like any other code.
Where you put the band depends on your data. The probability comes from the text you handed over. It doesn’t tell you how often Jev is right. Run the question against cases you have labeled first. thinkthen audit grades saved answers against your labels. Pick the band from that run.
Ten functions that pipe together
ThinkThen turns Jev’s three kinds of answer into ten functions. Each reads standard input and writes standard output. This pipeline keeps the buying inquiries, ranks them by how ready the buyer is, and picks a team for each:
buying="Is this a buying inquiry?"
ready="Is this buyer ready to pay now?"
team="Which team should take this?"
teams=(
enterprise
smb
)
cat <<'EOF' |
We need 200 seats next quarter. Please send a quote.
Please remove me from this list.
Our team of six wants to buy today. How do we pay?
Loved your talk at the conference last week.
EOF
thinkthen filter "$buying" --batch 1 |
thinkthen rank "$ready" |
thinkthen choose "$team" "${teams[@]}" --lines |
jq .{
"input": "Our team of six wants to buy today. How do we pay?",
"value": "smb"
}
{
"input": "We need 200 seats next quarter. Please send a quote.",
"value": "enterprise"
}The pipeline returns the original lines beside their answers. It doesn’t rewrite them.
Diogo Almeida of TypeSafe lists decisions Jev could make inside a coding agent, such as which tools to load for a step. Each one is a bounded question of this kind.
annotate answers a whole set of questions for every record. A set can mix decide, choose, score and tag questions in one file. Here is a set with one yes or no, one pick and one scale, saved as form.json:
{
"version": 1,
"questions": {
"steps": {
"decide": "Does the report give steps to reproduce?"
},
"area": {
"choose": "Which part of the app is this?",
"options": [
"export",
"login",
"billing"
]
},
"impact": {
"score": "How much does this block the user?",
"levels": [
"None.",
"Slows them.",
"Blocks work."
]
}
}
}Three bug reports go in, one per line. Each keeps its id and gains the three answers:
cat <<'EOF' |
{"id": "B-7", "body": "Steps: click Log in. Nobody gets in."}
{"id": "B-8", "body": "The Pay button on billing is too blue."}
{"id": "B-9", "body": "Steps: click Export. It is very slow."}
EOF
thinkthen annotate form.json \
--batch 1 \
--jsonl \
--field /body \
--jobs 8 |
jq .{
"id": "B-7",
"body": "Steps: click Log in. Nobody gets in.",
"steps": true,
"area": "login",
"impact": 1.98
}
{
"id": "B-8",
"body": "The Pay button on billing is too blue.",
"steps": false,
"area": "billing",
"impact": 0.09
}
{
"id": "B-9",
"body": "Steps: click Export. It is very slow.",
"steps": true,
"area": "export",
"impact": 1.04
}Where it breaks
Planted facts move the answer. Probe 06 asked jev-1.13.0 one question about twenty made-up messages: “The customer explicitly asks for money back.” Each message ran once clean and once with hostile text added. An order aimed at the model moved the probability of yes by 0.04 or less, across seventeen wordings. A false claim planted in the message moved it by as much as 0.57. Jev reads a planted claim and a true one the same way, because both look like evidence. A later review asked other questions of the same model. There an order moved a different question by 0.16 to 0.18. On that question, the order itself could count as evidence. A planted claim flipped a second question from 0.01 to 0.64. These numbers hold for those messages, questions and that model only. At the default bar of 0.5, a planted claim can flip an answer. Use a band, and send the middle to a person.
tag often misses part of the set. A song can have two lead singers, and tag must name every one to score. On Beatles Bench it names the whole set on 0.56 of songs. Its top label is a true lead on 0.80. The bench’s function table scores every function. Use tag to fill a queue a person reads. Don’t use it as a gate.
A threshold belongs to one model. ThinkThen speaks System One, the request format Jev answers. Any server that speaks System One can answer at another address. Its answers will differ, so tune the band again for each model.
It only answers. ThinkThen writes no text, holds no conversation and takes no action. The answer goes back to your code, and your rules decide what happens next.
What ships
ThinkThen ships 10 functions, 1 CLI and 24 bindings. Every binding calls the same Rust engine, so a question file reads the same way everywhere. In a database, a question sits in a WHERE clause like any other condition.