Ollama Local: Commands and Pitfalls to Avoid

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Ollama serves language models locally via an HTTP server and OpenAI-compatible APIs. Behind the ease of setup, stability depends on available memory, system environment variables, and the chosen output format. Here are the commands and settings to know to avoid slowdowns and unpleasant surprises.
macOS System Variables and Context: Settings That Block
On macOS, the Ollama desktop application is launched by the system and does not take into account the export lines from a .zshrc file. Thus, a configuration that seems correct in a terminal may not apply to the desktop application. To ensure that environment variables are recognized, they must be defined via launchd using the command "launchctl setenv." This method is often necessary for the context length or the model directory to be effectively modified. Additionally, the command "ollama ps" allows you to check the allocated context length, which may differ from the expected or requested value.
Monitor GPU/CPU and Memory with "ollama ps"
The command "ollama ps" displays the models currently in memory and includes a PROCESSOR column. If the displayed value is less than "100% GPU," it means that part of the model is using the CPU, which significantly slows down generation. This command also indicates the context actually allocated, useful for verifying the alignment between parameters and available resources.
Integrate Locally with OpenAI-Compatible Endpoints
Ollama sets up an HTTP server on port 11434 and offers the endpoints /api/chat and /api/embed. A /v1/ compatibility layer allows an existing OpenAI client to point to localhost without further modification. The service thus provides an OpenAI-compatible endpoint accessible from the local machine.
One-Command Startup, Structured Formats, and Customization
Running a model locally with Ollama requires just a single command, which downloads the necessary weights. For structured outputs, providing a JSON schema constrains the decoding to this form, ensuring systematic parsing, while the raw "json" format produces valid JSON without guarantees on the keys. Modelfiles allow you to save a base model with default parameters. Environment variables control the loading time of models and the number of models that can run in parallel. Finally, disk management commands are available to complement regular usage.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.