Brief IA

GPT-Live-1 API: $0.05/min and Customer Feedback

🤖 Models & LLM·Tom Levy·

GPT-Live-1 API: $0.05/min and Customer Feedback

GPT-Live-1 API: $0.05/min and Customer Feedback
Key Takeaways
1GPT‑Live‑1 is available in the API at $0.05 per minute for the voice layer
2Yelp and Speak report improvements in call management and reduced interruptions
3The model handles listening and speaking in a single block, delegates reasoning in the background, and offers more voices and customization options
💡Why it mattersGPT‑Live‑1 aims to make voice interactions more natural and adaptable for phone applications and services.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

A real-time voice model is now available as an API, priced at $0.05 per minute for the voice layer. Yelp and Speak report improvements in call management and interruptions, while developers gain levers for customizing style and voices, including in telephony.

Available at $0.05 per minute, with reported gains from Yelp and Speak

GPT‑Live‑1 is accessible through the API at a rate of $0.05 per minute for the front-end voice layer. Alex Levy, CTO of Yelp, states that integrating GPT‑Live‑1 into Yelp Host and Hatch has improved turn-taking and accuracy compared to the previous voice architecture. Yelp reports significant improvements in call management rates when Yelp Host handles reservations and food orders with this model, noting that callers are forming more complete and natural sentences. Speak reports a nearly 80% reduction in interruptions compared to turn-taking systems during initial evaluations. Andrew Hsu, co-founder and CTO of Speak, claims that the model brings a natural tutoring quality to live lessons. The model can be paired with an agent and a background model tailored to the product, to build scalable voice experiences according to tasks.

Full-duplex telephony for reservations and support

The model enables the establishment of voice agents capable of operating in full-duplex during phone calls, for tasks such as managing reservations or customer support. At Yelp, the integration into Yelp Host for handling reservation and order-related calls is accompanied, according to the company, by an increase in call handling rates. Yelp also reports progress in turn-taking and accuracy with this setup.

Customizable voices, accents, and conversational style

The range of real-time voices is expanding to include more accents, dialects, and languages, offering more choices for the sound signature of assistants. Developers can adjust the tone, pace, and style of exchanges via the system prompt. The API version emphasizes these steering and customization capabilities based on users, workflows, and objectives.

A single model for listening and speaking, with background delegations

GPT‑Live‑1 combines listening and speaking into a single model, simplifying the voice layer. This integration aims to avoid latency and fragile transitions between speech recognition, LLM, and speech synthesis by reasoning about incoming and outgoing audio jointly. The model can handle live interruptions and acknowledgments, then delegate deeper reasoning in the background, allowing the conversation to continue while the work is done. Delegations can rely on a text model like GPT‑6 Astra or a third-party model. In contrast, traditional architectures segment recognition, reasoning, and synthesis, requiring additional coordination on the developer's side, especially during pauses, interruptions, or changes in direction. The device listens and speaks simultaneously and extends delegation capabilities already highlighted with Codex and ChatGPT Work. GPT‑Live‑1 was initially introduced in ChatGPT.

Managed noise, long sessions, and suggested testing scenarios

The model better manages background noise and silence periods, without cutting off the conversation or verbalizing every step. Over time, it improves context retention and the quality of exchanges. It is suggested to experiment with natural conversations, including in noisy environments, to interrupt during responses to clarify a request, to try walking or in everyday noise, and to introduce hesitations or brief interactions with third parties before resuming the exchange.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.