OpenAI Unveils ChatGPT Images 2.0: Innovation and Challenges
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI Reinvents Image Generation
OpenAI recently announced the launch of ChatGPT Images 2.0, a next-generation image generation model that focuses on accuracy, usability, and complex visual tasks. This model stands out for its ability to combine text and images to create complex and aesthetic pages. OpenAI is thus redefining image generation, presenting it as a visual language rather than just a decorative creation process.
OpenAI describes this model by stating: "A good image does what a good sentence does: it selects, organizes, and reveals. It can explain a mechanism, create an atmosphere, test an idea, or make an argument."
Reflective Capabilities for Complex Workflows
In addition to its enhanced ability to blend text and graphics, the new model utilizes advanced reflective capabilities. It can generate multiple images per prompt with continuity between the results. This approach is possible because the model genuinely integrates reasoning into the image output.
For example, from a vague prompt like "Generate an infographic on activities to do considering tomorrow's weather in San Francisco," the AI will gather data on the weather and appropriate activities, then construct an image or set of images that correspond to the results.
According to OpenAI, "In this model, Images 2.0 acts more like a visual reflection partner, helping to take a project from a rough concept to a finished asset with much less work on your part."
Accuracy and Design Control
Many users have long struggled to convince ChatGPT to generate images in a specific aspect ratio. Often, the AI stubbornly produces what it wants. But now, with Images 2.0, the model supports "aspect ratios as wide as 3:1 and as tall as 1:3."
The model also supports high-fidelity outputs that produce (mostly) accurate object placement, detailed text rendering, and complex compositions. We will see if we can remove the word "mostly" from this sentence after the official product release.
Preview Testing
I had access to a preview the day before the release, and the model is impressive, for the most part. I provided a screenshot of the ZDNET homepage and a draft of the Images 2.0 press release.
Then, I asked: "Based on the content of the press release, generate a 16:9 infographic on the new image update and generate it using the ZDNET brand style as indicated in the ZDNET homepage document."
The model successfully created the infographic, but it could not reproduce the ZDNET logo. On its first attempt, it rendered the "Z" of ZDNET with a slight droop.
I tried a variety of requests to correct the logo, but Images 2.0 never managed to fix it. So, I started a new session, including the instruction: "Please ensure to accurately reproduce the ZDNET logo."
Issues with the Logo
It was at this point that things became strange. For its first attempt, the model retrieved a copy of the ZDNET logo from before our 2022 redesign, which is nowhere visible on our current homepage. Strangely, it rendered this old logo with the current color scheme. The model then pushed the logo and the infographic information off the left edge of the image, choosing a light blue color for "Images 2.0" that is not a ZDNET brand color.
I did my best to convince it to use the current logo, but even after asking not to look for an alternative logo, the issue remained unresolved.
Conclusion and Availability
I am testing a preliminary version of Images 2.0 and will return with a much more comprehensive test after the official product release. The new model is available today for all ChatGPT and Codex users. Advanced outputs and reflective capabilities are accessible to ChatGPT Plus, Pro, Business, and Enterprise users.
Currently, before the release, the new Images 2.0 model is only available on desktop, but OpenAI promises that these capabilities will also be present in the mobile version. Images are also available via the API using the gpt-image-2 model. API pricing varies based on the desired quality, complexity, and image resolution.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.