Local AI with Hermes: Uses, Costs, and Limitations on Mac Studio

A user has set up Hermes, a local AI agent, on a Mac Studio M5 Ultra with 256 GB of memory. Between useful automations and repeated failures, their experience sheds light on what these systems can already do, what they require in practice, and what remains out of reach.
Persistent Failures and Scripts Still in Development
The daily briefing set up failed several times before it started working consistently. According to the author, even powerful local models do not behave like all-capable assistants. A script intended to automate a series of laptop performance tests is still under development. Guiding Hermes and obtaining generated Python scripts has proven difficult, and the work continues. Today, benchmarks are still conducted manually, with a dozen tests repeated three times to calculate an average.
Concrete Uses: Briefing, Steam Sorting, and Sensitive Calculations
A morning briefing has been configured to analyze emails and calendars, report emergencies, and deliver the weather. For it to trigger at 7:30 AM, the machine must not be in sleep mode; after a rocky start, it now runs correctly. On the leisure side, Hermes was used to reorganize a Steam library of over 400 titles that the platform does not classify on its own. After installation, the agent detected most games via the local client and proposed organizational schemes. The user opted for a genre-based classification while keeping their custom categories. The operation required local permissions and the use of a revoked Steam web API key once the sorting was completed, which took just a few minutes. The agent also handled sensitive local tasks: calculations from financial documents and the creation of a comparative specifications table for a laptop whose information was under embargo, all without resorting to the cloud.
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Preparing and Operating Hermes: Heavy Models, Bot, and Scheduled Tasks
Hermes, an open-source self-hosted agent that is free when relying on local LLMs, was installed on a Mac Studio and quickly brought online. Remote control was enabled via a Telegram bot. The user selected, through Hermes' interface, the Qwen 3.8 Flash Next model, equipped with 125 billion parameters for about 105 GB. The number of available models and their specializations complicate the choice, with a general trend towards better performance for larger models. Local execution avoids token costs. For scheduled tasks, such as the briefing, the system must not be in sleep mode at the time of execution.
Privacy and Assumed Scope of Use
The ability to run powerful models locally addresses an initial reluctance to transmit personal data to cloud services. The author prefers an assistant that sends nothing to OpenAI, Google, Microsoft, or Anthropic and limits usage to tasks compatible with this framework. They do not consider entrusting it with writing, video editing, or visual creation, treating Hermes as a tool without polite interaction. They intend to remain cautious as they continue their trials, especially when information is under embargo and must stay on the machine.
Machines and Memory: From Mac Studio to RTX Spark PCs
The Mac Studio M5 Ultra used features 256 GB of unified memory, allowing for the loading of large models during tests conducted across multiple systems. Trials with smaller Qwen models are planned on a Mac Mini M6, a MacBook Air M5, an Asus TUF Gaming A14 with an AMD Strix Halo processor, and then on RTX Spark PCs. Apple highlights the efficiency of its desktop Macs for local AI. On the Windows side, a new wave of RTX Spark machines is announced with up to 128 GB of RAM for agent-based AI. Local AI and agents are presented as very trendy, with configurations reaching budgets of up to $12,000 for highly equipped stations.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.