
Google’s multimodal AI model for video
Gemini Omni is a new model family from Google that processes text, images, audio and video in a single prompt and uses this to generate video. Google unveiled the model on 19 May 2026 at the Google I/O developer conference. The first model available is called Gemini Omni Flash and has been live since launch day in the Gemini app, in Google Flow, and free of charge in YouTube Shorts Remix and YouTube Create. Clips are currently limited to ten seconds; this is a deliberate decision as part of the roll-out, not due to any technical limitation. This is relevant for marketing teams because quick, cost-effective video options are now readily available — albeit with clear restrictions on length and audio editing that one should be aware of.



