ThinkThen

Reference

How answers work

The model gives a probability for each choice. ThinkThen turns those probabilities into the answer it prints. Each function does this its own way.

Not sure

When decide or choose is not sure, it prints null and exits 3. When find --none finds nothing, it prints nothing and exits 3. A null never means a failure. A failure never prints null as its answer, and it exits with its own exit code. In annotate, a question that failed carries a failed marker instead of a value. Under exit 6, relate prints only the edges that succeeded.

An answer is not sure in three ways:

  • A decide probability falls inside a band.
  • A choose winner falls under its cut, or the top two options tie exactly.
  • find --none finds that nothing fits.

In a stream of records, each not-sure record prints null in its row. The run itself exits 0 when it finishes.

Bands and cuts

Only decide takes a band, written LOW:HIGH. A probability of yes at or above HIGH is yes. A probability under LOW is no. Anything between is not sure. Every other function that takes a threshold takes one number, a cut. The reference gives every boundary, and Set a cut or a band runs both on one message.

The refund request answers true and the thank-you note false. "I want to send this back." could mean an exchange or money back. It lands inside the band 0.2:0.8 as null.
question="Does the customer ask for a refund?"

cat <<'EOF' |
Please refund my order. It arrived broken.
Thanks for the quick help yesterday!
I want to send this back.
EOF
thinkthen decide "$question" \
  --batch 1 \
  --lines \
  --threshold 0.2:0.8 |
jq .
Output
{
  "input": "Please refund my order. It arrived broken.",
  "value": true
}
{
  "input": "Thanks for the quick help yesterday!",
  "value": false
}
{
  "input": "I want to send this back.",
  "value": null
}
exit 0

choose

choose takes a cut and never a band. When the winning option's probability falls under the cut, the answer is not sure. When the top two options tie exactly, the answer is also not sure, with or without a cut. The order of your options does not break a tie. choose never exits 1. The choose specification holds the rule.

Each of the first three messages names one team: billing, shipping, and account. The fourth names a parcel and a login. No team reaches 0.9, and it comes back null.
question="Which team owns this?"
teams=(
  --option "billing=Invoices, fees, and refunds."
  --option "shipping=Parcels and delivery."
  --option "account=Logins and passwords."
)

cat <<'EOF' |
Please refund the extra fee on my invoice.
My parcel went to the wrong address.
I cannot reset my password.
My parcel never came, and now I cannot log in to track it.
EOF
thinkthen choose "$question" "${teams[@]}" \
  --batch 1 \
  --lines \
  --threshold 0.9 |
jq -r .value
Output
billing
shipping
account
null
exit 0

find

find has no threshold. It reads every line together and prints the line with the highest probability. When two lines tie at the top, the first line wins. --none adds one more choice: nothing fits. When that choice leads, or ties for the lead, find prints nothing and exits 3.

No line says how to cancel. find prints nothing and exits 3.
question="Which line says how to cancel an order?"

cat <<'EOF' |
Returns need the original receipt.
Refunds are issued within 30 days of purchase.
Shipping is free on orders over $50.
Gift cards cannot be exchanged for cash.
EOF
thinkthen find "$question" --none
No outputexit 3: nothing fits, under --none

recognize

recognize keeps a name when its strength reaches --threshold. The cut defaults to 0.5 and must be above 0. Strength is the probability of the name's kind times the probability of its span, rounded to four places. It ranks names, and it is not a probability. --relation-threshold is a separate cut on each relation edge's probability, and it defaults to 0.5. recognize never exits 1 or 3. A text with no names is a good answer.

recognize finds three names. Each comes back with its kind and its strength.
person="PER=Part of a person's name."
org="ORG=Part of the name of an organization:"
org+=" a company, band, team, agency, government"
org+=" body, or media outlet."
place="LOC=Part of the name of a place: a country,"
place+=" region, city, or geographic feature."
other="MISC=Part of another named entity: a"
other+=" nationality, an event, a product, or the"
other+=" name of a creative work."
text="Maria Chen joined Northwind Freight in Chicago"
text+=" last spring."
printf '%s' "$text" |
thinkthen recognize \
  --kind "$person" \
  --kind "$org" \
  --kind "$place" \
  --kind "$other" |
jq -c '.entities[] | [.text, .kind, .strength]'
Output
["Maria Chen","PER",0.9987]
["Northwind Freight","ORG",0.997]
["Chicago","LOC",1.0]
exit 0

score

score has no threshold and no not-sure answer. You name 2 to 10 levels, lowest first. The number it prints is a position on those levels, counted from 0. ThinkThen multiplies each level's probability by its position, adds the results, and divides by the total of the probabilities it accepted. The number is a weighted average of the positions. A number between two positions means the probability is spread over more than one level. In a question file, levels may be a map from each level's name to its description. The model then reads the descriptions, and the positions stay the same. The score specification works one example through.

The address change lands near 0, the Friday deadline near 1, and the login outage at 2.
question="How urgent is this?"
levels=(
  "Routine."
  "Soon."
  "Immediate."
)

cat <<'EOF' |
Please update my mailing address when you can.
Can you send the signed contract by Friday?
Nobody can log in to the site right now.
EOF
thinkthen score "$question" "${levels[@]}" \
  --batch 1 --lines |
jq .
Output
{
  "input": "Please update my mailing address when you can.",
  "value": 0.06
}
{
  "input": "Can you send the signed contract by Friday?",
  "value": 0.99
}
{
  "input": "Nobody can log in to the site right now.",
  "value": 2.0
}
exit 0

rank

rank takes no threshold and prints every record, unless --top N asks for the first N. It sorts by the probability of yes, most likely first. With a saved score question, rank @FILE sorts by the score's position instead, highest first. Records with exactly equal values keep their input order. rank asks about each record on its own and sorts on this machine. It never asks the model to compare two records.