⚡
Brief IA
›

AI in Production: Investing in Dedicated Capacity?

🔬 Research·Tom Levy·

AI in Production: Investing in Dedicated Capacity?

AI in Production: Investing in Dedicated Capacity?
⚡
Key Takeaways
1The advancement of AI in production is leading companies to choose between pay-per-use and dedicated capacity.
2Owning AI capacity is only relevant beyond a certain level of sustained usage and with operational discipline.
3According to Deloitte, employee access to AI increased by 5% in 2025, and the share of companies with at least 40% of AI projects in production is expected to double within six months.
💡Why it matters — The right business model choice is crucial for cost control and large-scale value creation.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

With the industrialization of use cases, the cost of AI is no longer limited to token prices. Between pay-per-use and dedicated capacity, the choice depends on a sustained level of usage and rigorous management. Deloitte's figures indicate a recent increase in employee access and a likely acceleration in production deployment, making this decision even more immediate.

Making AI capacity productive requires continuous execution

Deciding to invest is only half the battle. Even when the economy favors ownership, capacity only creates value if workloads are quickly put into production and kept operational. Generating value requires much more than simply setting up infrastructure. An operational model must ensure the connection between technology, adoption, and business outcomes: this includes integrating users and workloads, governance and reviewing usage, as well as continuously identifying high-value use cases. It is also essential to monitor usage, detect unused capacity, and gradually integrate the most prioritized workloads. Without this discipline, the expected economic value may never materialize; with it, capacity becomes a productive, optimizable asset that generates measurable value.

Three checks before committing infrastructure budgets

Before allocating capital to dedicated capacity, three questions structure the decision: Is the demand already stable, predictable, and of sufficient volume to justify reserved capacity? At what level of usage does ownership become economically relevant? Can the organization maintain this productive capacity through adoption, governance, and continuous expansion of use cases? When usage is sustainable and sufficient to support productive capacity, the alternative then arises between continuing to buy on demand or investing in capacity that the company can optimize and control.

Cost shifts with usage, without a universal threshold

Owning capacity is not always the least expensive solution: it only makes sense if the organization can keep that capacity active. Each company has a threshold of sustained usage beyond which owning capacity may prove more cost-effective than purchasing on demand, without a universal value existing. This threshold varies depending on the models used, the distribution between input and output tokens, performance needs, how the system is designed, energy expenditures, and the chosen operational model. When usage reaches the appropriate level, owning capacity can lead to lower effective costs and better predictability, treating this asset as an infrastructure investment rather than a fluctuating monthly expense.

Varied workloads, divergent cost profiles

AI production today encompasses assistants, retrieval and knowledge systems, as well as agentic applications. These agents, whether deployed for customer service, IT, research, or processes, perform multi-step action sequences, leading to regular demands on models, data, and tools. A retrieval-based knowledge system processes more context per interaction than a simple assistant and presents a different cost profile; agentic workflows, on the other hand, combine repeated reasoning, retrieval, model calls, and tools. Generic cost references are therefore insufficient: it is necessary to model actual workloads, estimate demand, and size capacity accordingly. When multiple workloads share the same infrastructure, fixed costs can be amortized over more productive usage, improving the economics of ownership.

Adoption signals that bring economic arbitration closer

Field signals are accumulating. Deloitte observes a 5% increase in employee access to AI by 2025 and anticipates a doubling, within six months, of the proportion of companies with at least 40% of projects in production. The shift towards active workloads is underway, along with a new economy: pay-per-use remains flexible and low-commitment, but the key question becomes predictability and sustainability at scale. The arbitration does not dogmatically oppose cloud and on-premises: it is decided workload by workload, projecting demand over 12 to 18 months and usage frequency. The most successful organizations look beyond token prices and the latest model, know how to recognize when recurring demand necessitates a different economic model, and have the discipline to make capacity productive. When these conditions are met, AI ceases to be an expense and becomes an asset. Conversely, focusing the discussion solely on unit prices and consistently aiming for the most efficient model is not always necessary and can shift the problem towards a volatile monthly expense, influenced by usage, workloads, and model requirements. The real challenge is to operate AI economically, predictably, and at a sustained scale.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.