ThinkThen

for anyone who learns from talks

Search YouTube transcripts by meaning

Ask a question of a talk, and find the passages that answer it. paste joins the caption lines into passages. rank puts the best answers first. filter keeps the passages that answer yes.

fetch.sh downloads the automatic captions of a YouTube talk with yt-dlp. Change the id to search another talk.
id=YbzrlpAyCV4
yt-dlp \
  --skip-download \
  --write-auto-subs \
  --sub-langs en \
  --sub-format vtt \
  -o talk \
  "https://youtu.be/$id"
awk -f vtt-to-lines.awk talk.en.vtt > transcript.txt
vtt-to-lines.awk writes one timed line per caption. Automatic captions repeat each line as it scrolls, so it keeps only the lines that carry new words.
/-->/ {
  split($1, t, /[:.]/)
  if (t[1] > 0)
    stamp = sprintf("%d:%02d:%02d", t[1], t[2], t[3])
  else
    stamp = sprintf("%d:%02d", t[2], t[3])
  next
}
/<c>/ {
  gsub(/<[^>]*>/, "")
  gsub(/&gt;/, ">")
  gsub(/&lt;/, "<")
  gsub(/&#39;/, "\047")
  gsub(/&quot;/, "\"")
  gsub(/&amp;/, "\\&")
  if (length($0)) print "[" stamp "] " $0
}
measured.txt records the full experiment. ThinkThen searched three whole talks, 292 passages, in under half a second per question. Each question cost about a quarter of a cent.
Measured 2026-10-02 with thinkthen 0.1.0 and jev-1.13.0.
Three whole talks, 292 passages of 20 caption lines each.
Each run sent every answer live with --refresh-cache.
Cost is input tokens at $0.042 per million.
The rows copy the run's cost.txt, with narrower spacing.

rank-excited.txt    0.4 s  61031 input tokens  $0.00256
rank-problem.txt    0.4 s  62199 input tokens  $0.00261
filter-problem.txt  0.3 s  62199 input tokens  $0.00261
filter-problem.txt kept 4 of the 292 passages.
transcript.txt holds 200 timed lines from the talk Introducing ThinkThen, in the form fetch.sh writes. They cover the first four minutes and the minutes from 15:45 on. The speaker corrected these captions by hand.
wc -l < transcript.txt
head -3 transcript.txt
Output
200
[0:00] Hello, my name is Ian Maurer. I'm the CTO
[0:02] of GenomOncology and I'm here to
[0:03] introduce ThinkThen, my new library for
exit 0
awk drops the time from every line but the first of each ten. paste joins each ten lines into one passage of about 20 seconds. The 200 lines make 20 passages.
awk 'NR % 10 != 1 { sub(/^\[[0-9:]+\] /, "") } 1' \
  transcript.txt |
paste -d' ' - - - - - - - - - - > passages.txt
wc -l < passages.txt
cut -c1-60 passages.txt | head -2
Output
20
[0:00] Hello, my name is Ian Maurer. I'm the CTO of GenomOnc
[0:22] don't know that Abbey Road is an album or the song Oc
exit 0
rank judges all 20 passages and prints the three most likely to answer yes. The top passage says a large language model could always do this work, but slowly and at a higher cost. The next two say it is zero shot and needs no labels or training.
asks="Does this passage explain"
question="$asks why people are excited about Jev?"

thinkthen rank "$question" --top 3 < passages.txt |
cut -c1-60
Output
[17:59] looking at you know you know did did Jenkins fail di
[1:59] don't have to do what you'd have to do in the past wi
[1:36] performant and very easy to use. And it's zero shot m
exit 0
filter keeps a passage when its answer is yes at the default cut of 0.5. This question asks for more, and one passage clears the cut. Use rank to explore and filter to keep.
asks="Does this passage explain"
question="$asks what problem Jev solves that LLMs do not?"

thinkthen filter "$question" < passages.txt |
cut -c1-60
Output
[17:59] looking at you know you know did did Jenkins fail di
exit 0

The functions it uses