ThinkThen

Reference

annotate's edge cases

annotate answers a saved set of questions for every record. Where each answer lands depends on the record. This page uses one question set throughout.

form.json, three questions
{
  "version": 1,
  "questions": {
    "steps": {
      "decide": "Does the report give steps to reproduce?"
    },
    "area": {
      "choose": "Which part of the app is this?",
      "options": [
        "export",
        "login",
        "billing"
      ]
    },
    "impact": {
      "score": "How much does this block the user?",
      "levels": [
        "None.",
        "Slows them.",
        "Blocks work."
      ]
    }
  }
}

Flat fields

An object record gains one top-level field per question, named for the question. Every field it already held rides through unchanged. CSV and TSV rows are objects of strings. They stay flat too.

One JSON document goes in. The same document comes back with three answers added: steps, area, and impact. The nested report rides through unchanged.
cat <<'EOF' |
{
  "id": "B-7",
  "report": {
    "page": "/login",
    "body": "Steps: click Log in. Nobody gets in."
  }
}
EOF
thinkthen annotate form.json \
  --field /report/body |
jq .
Output
{
  "id": "B-7",
  "report": {
    "page": "/login",
    "body": "Steps: click Log in. Nobody gets in."
  },
  "steps": true,
  "area": "login",
  "impact": 1.98
}
exit 0

The wrapper

A record that is not an object has no field to add to. In a stream of records, a text line, a JSON scalar or a JSON array comes back as {"input":…,"value":…}. The record sits under input, and the named answers sit under value. On one document that is not an object, annotate prints the named answers alone.

A line is text, not an object. Its answers come back under value, and the line itself under input.
cat <<'EOF' |
Steps: click Log in. Nobody gets in.
EOF
thinkthen annotate form.json --lines |
jq .
Output
{
  "input": "Steps: click Log in. Nobody gets in.",
  "value": {
    "steps": true,
    "area": "login",
    "impact": 1.98
  }
}
exit 0

A name clash

A question name the record already holds would overwrite the record's own value. annotate refuses that record at exit 2, before any request for it.

The record already holds steps. annotate refuses it at exit 2 and sends nothing for it.
cat <<'EOF' |
{
  "id": "B-7",
  "steps": "none given",
  "report": {
    "page": "/login",
    "body": "Steps: click Log in. Nobody gets in."
  }
}
EOF
thinkthen annotate form.json \
  --field /report/body
Output
thinkthen: the record already holds `steps`, so that question cannot be appended
exit 2: usage or input error

--plan checks records in order and stops at the first clash. It sends nothing. A plan finds a clash at no cost. It names only the first clash it meets.

The second record holds area. The plan refuses it at exit 2 and sends nothing.
cat <<'EOF' |
{"id": "B-7", "body": "Steps: click Log in. Nobody gets in."}
{"id": "B-8", "body": "The Pay button is too blue.", "area": "billing"}
EOF
thinkthen annotate form.json \
  --jsonl \
  --field /body \
  --plan
Output
thinkthen: the record already holds `area`, so that question cannot be appended
exit 2: usage or input error

--details keeps the record under input

--details prints each record as input, value, answers and meta. The original record stays whole under input, and the answers sit under value. The two never share a level, so no name can clash. Use it when the records may hold any field name.

Under --details the record keeps its own steps under input. The answer named steps sits under value.
cat <<'EOF' |
{
  "id": "B-7",
  "steps": "none given",
  "report": {
    "page": "/login",
    "body": "Steps: click Log in. Nobody gets in."
  }
}
EOF
thinkthen annotate form.json \
  --field /report/body \
  --details |
jq '{input, value}'
Output
{
  "input": {
    "id": "B-7",
    "steps": "none given",
    "report": {
      "page": "/login",
      "body": "Steps: click Log in. Nobody gets in."
    }
  },
  "value": {
    "steps": true,
    "area": "login",
    "impact": 1.98
  }
}
exit 0

A missing pointer and exit 7

A question may read one part of each record with on, a JSON Pointer. Here each question reads /body.

on-body.json, the same questions on /body
{
  "version": 1,
  "questions": {
    "steps": {
      "decide": "Does the report give steps to reproduce?",
      "on": "/body"
    },
    "area": {
      "choose": "Which part of the app is this?",
      "options": [
        "export",
        "login",
        "billing"
      ],
      "on": "/body"
    },
    "impact": {
      "score": "How much does this block the user?",
      "levels": [
        "None.",
        "Slows them.",
        "Blocks work."
      ],
      "on": "/body"
    }
  }
}

A record without that part stops the run at exit 2. --on-error continue skips it instead. It needs --jsonl, --details and --batch 1, and it refuses --plan. It recovers only a missing pointer. The skipped record gets one thinkthen.record-error/1 row in its place, and no request goes out for it. Standard error counts the skipped records. A run that skipped any record exits 7.

B-8 has no /body. It gets one error row in its place. B-7 and B-9 get their answers. The run exits 7, and the test line checks it.
cat <<'EOF' |
{"id": "B-7", "body": "Steps: click Log in. Nobody gets in."}
{"id": "B-8", "text": "The Pay button on billing is too blue."}
{"id": "B-9", "body": "Steps: click Export. It is very slow."}
EOF
thinkthen annotate on-body.json \
  --jsonl \
  --details \
  --batch 1 \
  --on-error continue \
  > rows.jsonl
skipped_code=$?
jq -c 'if .schema == "thinkthen.record-error/1"
  then .
  else {id: .input.id, value}
  end' rows.jsonl
test "$skipped_code" = 7
Output
thinkthen: 1 record skipped
{"id":"B-7","value":{"steps":true,"area":"login","impact":1.98}}
{"schema":"thinkthen.record-error/1","at":2,"failure":{"kind":"usage","cause":"missing_pointer","pointer":"/body"}}
{"id":"B-9","value":{"steps":true,"area":"export","impact":1.04}}
exit 0

Failed questions and exit 6

When the backend cannot answer one question, that question fails and the others keep their answers. The failed question's field holds {"failed":{"kind":"backend","cause":…}}, never null. null means not sure. The run finishes and exits 6. When a reply holds no usable answer at all, the run stops at exit 4. It prints none of that record's answers, even those another on group answered. When a run has both failed questions and skipped records, exit 7 wins. A failure is never saved. Under a cache, a later run asks only the failed questions again.