⚡
Brief IA
›

Ollama Local: Commands and Pitfalls to Avoid

🤖 Models & LLM·Tom Levy·

Ollama Local: Commands and Pitfalls to Avoid

Ollama Local: Commands and Pitfalls to Avoid
⚡
Key Takeaways
1"ollama ps" signals any overflow to the CPU and displays the allocated context
2On macOS, setting variables via "launchctl setenv" avoids ignored settings
3Endpoints /api/chat, /api/embed, and layer /v1/ to switch an OpenAI client
💡Why it matters — these settings and commands determine performance, parameter stability, and the integration of Ollama locally.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Ollama serves language models locally via an HTTP server and OpenAI-compatible APIs. Behind the ease of setup, stability depends on available memory, system environment variables, and the chosen output format. Here are the commands and settings to know to avoid slowdowns and unpleasant surprises.

macOS System Variables and Context: Settings That Block

On macOS, the Ollama desktop application is launched by the system and does not take into account the export lines from a .zshrc file. Thus, a configuration that seems correct in a terminal may not apply to the desktop application. To ensure that environment variables are recognized, they must be defined via launchd using the command "launchctl setenv." This method is often necessary for the context length or the model directory to be effectively modified. Additionally, the command "ollama ps" allows you to check the allocated context length, which may differ from the expected or requested value.

Monitor GPU/CPU and Memory with "ollama ps"

The command "ollama ps" displays the models currently in memory and includes a PROCESSOR column. If the displayed value is less than "100% GPU," it means that part of the model is using the CPU, which significantly slows down generation. This command also indicates the context actually allocated, useful for verifying the alignment between parameters and available resources.

Integrate Locally with OpenAI-Compatible Endpoints

Ollama sets up an HTTP server on port 11434 and offers the endpoints /api/chat and /api/embed. A /v1/ compatibility layer allows an existing OpenAI client to point to localhost without further modification. The service thus provides an OpenAI-compatible endpoint accessible from the local machine.

One-Command Startup, Structured Formats, and Customization

Running a model locally with Ollama requires just a single command, which downloads the necessary weights. For structured outputs, providing a JSON schema constrains the decoding to this form, ensuring systematic parsing, while the raw "json" format produces valid JSON without guarantees on the keys. Modelfiles allow you to save a base model with default parameters. Environment variables control the loading time of models and the number of models that can run in parallel. Finally, disk management commands are available to complement regular usage.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.