Ollama, the popular open-source runtime for running large language models locally, has released version 0.32.9. The headline feature is support for Nemotron 3.5 Lightning, a new model built on the Nemotron 3 architecture. This release also brings significant upgrades to tool-call parsing, making it easier for developers to integrate function calling into their local AI applications. Streaming inference has been improved as well, which should reduce latency for real-time interactions. For developers and teams running models on their own infrastructure, this update enhances both capability and usability, keeping Ollama competitive with cloud-based LLM services. The new model support also opens up more options for experimentation with different model families without needing to switch tools.
Ollama v0.32.9 introduces Nemotron 3.5 Lightning support, improved tool-call parsing, and streaming inference upgrades for local LLM deployments.