Beatles Bench · What Jev knows
Retrieval-augmented decisions (RAD)
Giving Jev context improves accuracy. Look up the record, put it in front of the question, and ask. We call it retrieval-augmented decisions. ThinkThen sends this context as state.
The song catalog covers 1,075 of the bench's questions. From memory, Jev gets 68% of them right. With the catalog in the text, the bench's report estimates about 97%. Context makes each request longer, and a longer request costs more.
The numbers come from the bench's record for this slide.
Run it
From memory, with the band 0.1:0.9, Jev still says yes. A Day in the Life is not on Abbey Road.
is_a_song="The text is the title of a song by the Beatles."
on_abbey_road="It appears on the album Abbey Road."
question="$is_a_song $on_abbey_road"
printf '%s' "A Day in the Life" |
thinkthen decide "$question" \
--threshold 0.1:0.9 \
--replay recordingOutput
true
exit 0: yes
Add context
With the catalog entry, the same band gives no.
about_an_entry="The text gives a catalog entry"
about_an_entry+=" and then names a song by the Beatles."
on_abbey_road="It appears on the album Abbey Road."
question="$about_an_entry $on_abbey_road"
album="Sgt. Pepper's Lonely Hearts Club Band (1967-05-26)"
entry="A Day in the Life (lead: Lennon;"
entry+=" written: Lennon–McCartney; 5:38;"
entry+=" released 1967-05-26; first album:"
entry+=" Sgt. Pepper's Lonely Hearts Club Band)"
printf 'Catalog:\n%s\n%s\nText: %s' \
"$album" \
"$entry" \
"A Day in the Life" |
thinkthen decide "$question" \
--threshold 0.1:0.9 \
--replay recordingOutput
false
exit 1: no
The lesson
No sensible bar fixes a miss this sure. With the catalog entry, the answer turns right.
When memory fails, put the facts in the text.
On GitHub: github.com/botassembly/beatles-bench/tree/main/examples/decide