ThinkThen

Functional patterns

A ThinkThen question works like a function. Bind the question once. Then apply it to text, a list, a column or a stream.

A judge is a question with no input

Call tt.decide with a question and no text. It returns a judge. A judge is a plain Python callable. functools.partial and toolz.pipe take it like any function.

Hand the judge the whole list. map(judge, xs) calls the judge once per item and sends one request per item. judge(xs) hands over the whole list. ThinkThen packs it into as few requests as THINKTHEN_BATCH and THINKTHEN_MAX_REQUEST_BYTES allow. tt.plan gives the count before anything is sent.

import thinkthen as tt

question = "Is this a complaint?"
reviews = [
    "Arrived a day early. Thank you!",
    "The zipper broke the first time I used it.",
    "Does this come in blue?",
    "The strap snapped on day two.",
]
is_complaint = tt.decide(question)

whole_list = tt.plan(is_complaint, reviews)
one_by_one = [tt.plan(is_complaint, [r]) for r in reviews]
assert whole_list["requests"] == 1
assert sum(p["requests"] for p in one_by_one) == 4

complaints = is_complaint(reviews).value
assert complaints == [False, True, False, True]

A stream stops when you stop

Give a judge a generator, and it returns a lazy stream. The stream reads ahead a little and packs what it read. When the reader stops, ThinkThen stops sending. Here the reader takes the first complaint and closes the stream.

import thinkthen as tt

question = "Is this a complaint?"
reviews = [
    "Arrived a day early. Thank you!",
    "The zipper broke the first time I used it.",
    "Does this come in blue?",
    "The strap snapped on day two.",
]
keep_complaints = tt.filter(question)

lines = (review for review in reviews)
with keep_complaints(lines) as complaints:
    first_complaint = next(complaints)
assert first_complaint == reviews[1]

A slow source can leave a request half full while it waits. Build the judge with batch=1 for a live source.

Count every call

A judge built with tally= adds the facts of each call it makes to the tally. This judge read four reviews, then two.

import thinkthen as tt

question = "Is this a complaint?"
reviews = [
    "Arrived a day early. Thank you!",
    "The zipper broke the first time I used it.",
    "Does this come in blue?",
    "The strap snapped on day two.",
]
tally = tt.Tally()
is_complaint = tt.decide(question, tally=tally)

complaints = is_complaint(reviews).value
first_two = is_complaint(reviews[:2]).value
assert complaints == [False, True, False, True]
assert first_two == [False, True]
assert tally.facts["records"] == 6

R

tt_decide with no input returns a function. Apply it inside mutate. Each group sends one packed call.

library(dplyr)
library(thinkthen)

is_refund <- tt_decide(
  "Does the customer ask for a refund?",
  threshold = "0.2:0.8"
)
messages <- tibble(
  shop = c("north", "south"),
  body = c(
    "Please refund my order. It arrived broken.",
    "I want to send this back."
  )
)

answered <- messages |>
  group_by(shop) |>
  mutate(is_refund = is_refund(body)$value) |>
  ungroup()
stopifnot(identical(answered$is_refund, c(TRUE, NA)))

The shell

The pipe is the same idea. filter reads a stream of lines and keeps the ones whose answer is yes.

filter keeps the lines whose answer is yes, in order.
question="Is this a complaint?"

cat <<'EOF' |
Arrived a day early. Thank you!
The zipper broke the first time I used it.
Does this come in blue?
The strap snapped on day two.
EOF
thinkthen filter "$question" --batch 1
Output
The zipper broke the first time I used it.
The strap snapped on day two.
exit 0: the run finished

More

The Python README covers judges over Polars and pandas columns, deadlines and cancel tokens. The R README covers dbplyr. The Rust Polars README covers lazy expressions. Caching and replay shows why a second run sends nothing.