Google's Gemini Omni: Revolution or Illusion for AI Video?

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Gemini Omni: Google's Any-to-Any Model
Gemini Omni is Google's first "any-to-any" model, capable of simultaneously understanding and generating text, images, audio, and video. This innovation allows for the processing of these different media types without the need for intermediary models, thereby simplifying complex processes. With this approach, it is possible to create videos from simple text descriptions or images, and to edit existing sequences with great precision.
The model stands out for its ability to integrate multiple types of media into a single workflow, which was previously impossible without using multiple distinct tools. This opens up new possibilities for content creators who can now generate complex and dynamic videos with unprecedented ease.
Interface and Generation Features
During our tests, the public application of Gemini proved to be intuitive and efficient. It allows users to choose the aspect ratio of the clip, whether landscape or portrait, before starting the generation process. This avoids approximations and ensures an optimal output. The interface also offers predefined templates to instantly apply specific graphic styles, making video customization easier.
This simplified user interface is designed to be accessible even to users without prior video editing experience, making the technology available to a broader audience. The predefined templates also save time by applying consistent visual styles with just a few clicks.
Editing Existing Videos
One of the main strengths of Gemini Omni is its ability to modify existing videos using natural language. Users simply describe the desired changes, and the model takes care of the rest, preserving the visual coherence of the sequence. For example, replacing a visual element in an existing video by providing a reference image is particularly useful for teams looking to integrate a product into a real scene. However, during our tests, an issue arose: Omni understood that all taxis should be replaced with a luxury car, which was not our initial intention.
In another example, we transformed a sunny scene in New York into a stormy one. Although Gemini generally adhered to the instructions given in the prompt, the visual quality of some elements left much to be desired. The generated lightning lacked realism, and the animation of the water appeared artificial.
This ability to modify existing videos without having to completely rebuild the sequence is a significant time-saver for content creators, but it still requires improvements to achieve an optimal level of realism.
Restrictions in Europe and Alternatives
It is important to note that within the European Economic Area, the audio and video input features of Gemini are currently restricted for regulatory reasons related to GDPR and the AI Act. However, these limitations can be partially circumvented by using Google Flow, accessible via the Google AI Pro subscription. Google Flow uses a credit system, allowing users to manage their resource consumption.
This geographical restriction poses an obstacle for European users who wish to fully leverage the capabilities of Gemini Omni. Google Flow offers a viable alternative, but it requires a paid subscription, which may deter some users.
Audio and Video Generation from Photos
Gemini Omni also stands out for its ability to generate sound synchronized with on-screen action. During our tests, while the requested sound effects were present, their synchronization left much to be desired. Additionally, the generation of videos from photos is a strong point of the model. We provided images of a location to generate a 360° video, but the final output failed to achieve the requested full rotation, settling for a partial panorama.
Beyond animating locations, this ability to generate videos from still images finds particularly relevant applications in advertising. For example, we tested creating a video from a sports shoe. The result is of high quality, although the transitions could be smoother. The shoe, however, is completely faithful to the provided image.
Creation of Digital Doubles
Another interesting feature of Gemini Omni is the creation of digital doubles, AI clones of your face and voice, usable in all generated videos. To create this digital double, the user must scan their face from multiple angles using the front camera, then read a few sentences aloud to model their voice. This feature is not yet available in the EEA, the UK, and Switzerland.
The creation of this digital clone is done directly from the Gemini app. Once generated, this avatar is linked to the user's Google account in the form of a tag (for example, @name). It then becomes extremely simple to integrate it into any video by simply mentioning this tag in the generation prompt.
This technology paves the way for new forms of personalization in video content, allowing users to digitally integrate themselves into various scenarios.
Series Consistency and Editing Tools
To ensure graphic consistency across a series of clips, Google Flow, an AI production tool, can be used. It allows for the generation of series clips with recurring characters. However, there are still inconsistencies in the backgrounds if multiple videos are generated in the same location, as the AI may not have a clear notion of this. Flow directly integrates an editing tool where users can insert the various generated clips and modify them directly with the agent. For now, it is only possible to edit one video at a time, which must be carefully selected.
In this example, we used Google Flow, an AI production tool developed by Google Labs, which utilizes Omni. Google Flow is available for paying subscribers and uses a credit system. First, a clean reference image of the character must be generated with Nano Banana. To ensure the generated character is effective, use a solid background (preferably white), request front and profile photos to improve the final video quality, and ensure no other person or face appears in the image.
For each clip, recall the character via @character (for example, @sophie in our case) in the prompt, with the Omni model activated. Then, drag the product. The advantage of Google Flow is that you can choose the duration, model, number of versions you want, and the desired orientation.
Be careful to always drag the character and the product in each message on Flow, otherwise the tool may use a different character. There are always inconsistencies in the backgrounds if multiple videos are generated in the same location, as the AI may not have a clear notion of this. A good idea is to create a storyboard with Nano Banana to generate consistent locations and characters throughout the video.
Digital Watermark and Pricing
All videos produced by Gemini Omni now incorporate Google's DeepMind SynthID technology, a digital watermark injected directly into the pixels of each image during generation. This watermark remains detectable even after modifications such as compression or cropping. In terms of pricing, three subscription levels are offered, ranging from the free plan to Google AI Ultra at €99.99/month, providing extended quotas and priority access to the latest models.
Gemini Omni represents a significant advancement in AI video creation, simplifying processes that previously required multiple distinct tools. However, geographical restrictions limit its potential in Europe. Users must weigh the cost of subscriptions and geographical limitations against the benefits offered by this cutting-edge technology.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.