Brief IA

GPT-5.4 Thinking: OpenAI Innovates but Sometimes Confuses

🤖 Models & LLM·Tom Levy·

GPT-5.4 Thinking: OpenAI Innovates but Sometimes Confuses

GPT-5.4 Thinking: OpenAI Innovates but Sometimes Confuses
Key Takeaways
1GPT-5.4 Thinking from OpenAI offers deeper analysis than its predecessors.
2The model sometimes answers unasked questions, which can confuse users.
3The quality of image generation and text formatting still lags behind.
💡Why it mattersGPT-5.4 Thinking highlights the current advancements and limitations of AI in handling complex tasks, influencing its professional adoption.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

OpenAI Unveils GPT-5.4 Thinking: A Major Advancement

OpenAI has recently launched GPT-5.4 Thinking, a version that marks a turning point in the evolution of artificial intelligence models. Unlike previous updates, this version does not merely offer minor improvements. In fact, the company has chosen to jump directly from version 5.2 to 5.4, highlighting a significant advancement. This model is specifically designed to handle complex cognitive tasks and is available for the Codex programming tool, the API, as well as for paid subscribers of ChatGPT. To evaluate its capabilities, I used the ChatGPT Plus plan, offered at $20 per month.

This new version, GPT-5.4 Thinking, is not simply an incremental update. It has been designed to provide deeper thinking capabilities, allowing it to tackle more complex challenges. Available for the Codex programming tool, the API, and the paid plans of ChatGPT, it stands out due to its advanced cognitive preparation. To test this model, I subscribed to the ChatGPT Plus plan at $20 per month.

A Demanding Test for GPT-5.4

Testing GPT-5.4 Thinking required a different approach than previous versions. Typically, my tests consist of a series of short and varied prompts. However, this model demands more elaborate scenarios and deeper challenges. The generated responses were often too detailed to be directly integrated into this article. Therefore, I chose to provide links to the complete test sessions, allowing readers to explore the responses in depth.

In general, when I test a version of ChatGPT, I subject it to a series of varied tests. Some are quick, while others are a bit more detailed. The prompts are usually short. The responses often lend themselves to being included in an article. However, this thinking model required deeper dives, with more comprehensive challenges. Thus, not only were the prompts more involved, but the responses were far too lengthy to be integrated into the article. Instead, I will provide links to each test session. By following the links, you can see the complete response in depth. Generally, a shared transcript opens at the end, so scroll up to get the full content of that discussion.

The Strengths and Weaknesses of the Model

Strengths

GPT-5.4 Thinking stands out for the quality of its textual responses. During testing, the model demonstrated impressive reasoning ability, addressing the posed challenges thoughtfully and without glaring errors. Each response provided constructive added value, which is an undeniable asset for users seeking precise and well-argued solutions.

Limitations

However, not everything is perfect. The model sometimes responded to questions different from those posed, which can be frustrating. Additionally, the quality of image generation and text formatting leaves something to be desired. The produced images do not meet expectations, and the formatting of texts, often in the form of long numbered lists, can seem awkward.

Testing Experiences: An Aircraft Carrier in the Sky

To evaluate the image generation capabilities of GPT-5.4 Thinking, I began with a visual challenge: creating an image of a flying aircraft carrier supported by turboprop engines. This test continued my experiments with other AIs, which often failed to position the engines correctly.

The result obtained with GPT-5.4 Thinking showed the same errors as its predecessors, with propellers oriented backward. However, the model provided detailed explanations about the design of such a craft, highlighting the technical constraints and potential tactical advantages.

I started with an image generation challenge. The initial prompt was: "Create an image of a flying aircraft carrier in the sky, supported by four upward-facing turboprop engines in round fan housings, carrying a squadron of fighters on its deck." I began with this because previous image generation tests across several AIs had not done well. They almost always oriented the engines toward the back of the aircraft carrier. Gemini Nano Banana 2 strangely placed the engines at the front, with the carrier moving in a forward thrust. Sometimes, we just don't want to know.

Anyway, right from the start, with the model set to GPT-5.4 Thinking, ChatGPT returned this image. As you can see, it has the same problem. Although if you look closely, the propellers are oriented toward the back of the plane, and there are visual thrust beams pulling downward. You win some, you lose some.

But then, I had an idea. Since this is the thinking model, what would happen if I asked it to design a helicopter? What would it propose? I specified the characteristics of the craft and then added these instructions: "Design such a vehicle, particularly explaining its structure and how it will be kept in the air, as well as any constraints or issues, along with any tactical advantages." I received a long, well-thought-out response. I particularly liked the section where it explained why "four downward-facing turboprop engines are a weak solution." It stated that they look dramatic, but it laid out a series of solid engineering reasons why it's a bad idea from an aeronautical construction standpoint.

Conclusion: A Promising but Imperfect Model

In conclusion, GPT-5.4 Thinking represents a notable advancement in the field of artificial intelligence, particularly for tasks requiring deep thinking. However, improvements are needed, especially regarding image generation and the clarity of responses. Despite these limitations, this model remains a powerful tool for users looking to solve complex problems.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.