Gemini Omni lets you re-edit a video by talking to it
Gemini Omni is Google DeepMind's video generation and editing model, the one taking over from Veo inside the Gemini app. Feed it text, photos, sound or existing footage and it builds a coherent clip, then reworks it line by line, like a conversation. Access runs through the Google AI Plus, Pro or Ultra plans, with a free rollout on YouTube Shorts and a per-second API.
- Video edits typed as plain sentences, no timeline
- Takes text, photos, audio and footage together
- Characters, lighting and props survive each edit
- Renders sharp, readable on-screen text
- Draws on Gemini's grasp of physics and history
- Paid Google AI plan required in the Gemini app
- API clips currently capped around ten seconds
- Adults only, with some country-level restrictions
Build a clip, then fix it by chatting
Gemini Omni generates video from any mix of text, images, audio and existing footage, then applies your corrections in plain language while keeping the rest of the scene intact. One instruction changes a detail, no full regeneration required.
Here is how it plays out: you drop in a product photo, type 'circle the camera around it like a drone', and the aerial shot lands. Then 'swap the background for a beach at sunset, keep everything else' moves the scenery without touching the object. (A practical habit worth stealing: spelling out what must stay put is half the craft.)
One rare bonus, on-screen text comes out crisp and legible, and the model leans on Gemini's knowledge of physics, history and science to keep scenes believable.
- Background swaps behind a subject
- Wardrobe and style changes mid-shot
- Object or character replacement
- Lighting adjusted through a written note
- Stabilization for shaky phone footage
Gemini Omni or Veo, two engines for two jobs
Inside the Gemini app, Omni replaces Veo, while Veo 3.1 stays available through the API for scene extension and last-frame control. Google itself now points to Omni as the default pick for video generation.
The split comes down to workflow. Omni remembers your edits from one turn to the next, where Veo shines on first-pass quality and higher resolutions. Both live in the same API, so starting on one and finishing on the other takes no migration at all.
What access costs, from YouTube to the API
Access to Gemini Omni depends on the door you pick. The Gemini app reserves it for Google AI Plus, Pro and Ultra subscribers, YouTube Shorts is rolling it out for free, Google Flow runs on plan-based credits, and the API bills per second of output.
On the API side, a 10-second clip, the current ceiling, runs a bit over $1 in output cost, with no free tier and no batch discount. These figures come from a public preview, so let Google's own pages have the last word before you budget anything.
| Access point | Best for | Billing |
|---|---|---|
| Gemini app | Google AI Plus, Pro or Ultra subscribers | Included in the plan |
| YouTube Shorts | Casual video makers | Free, gradual rollout |
| Google Flow | Production teams | Credits tied to the plan |
| API (gemini-omni-flash-preview) | Developers | $0.10 per second of video |
Frequently asked questions
Is Gemini Omni free?
Not in the Gemini app, where a Google AI Plus, Pro or Ultra plan is required and you must be over 18. Google is, however, rolling out Omni Flash at no cost inside YouTube Shorts, and developers pay the API per second of generated video. The answer hangs on which door you walk through.
Does Gemini Omni replace Veo?
Within the Gemini app, yes: Google states that Omni takes Veo's place there. Veo 3.1 remains alive on the API side, where it keeps distinct abilities like scene extension and last-frame control. In practice the two models complement each other rather than cancel each other out.
Is my footage used to train the model?
Yes for the API preview: Google flags that content sent to Omni Flash helps improve its products, even on the paid tier, unlike most other billed Gemini models. Generated clips also carry a SynthID watermark, an invisible marker identifying them as AI-made.
Can I make a video avatar of myself?
A personal avatar that looks and sounds like you is part of the Gemini Omni rollout. To curb deepfakes, Google asks you to record yourself reading numbers aloud before generating your double. Once verified, there is no re-uploading reference photos for every new clip.
Verdict: When a full-resolution hero shot on the very first try is the goal, Veo 3.1 pairs with Omni in the same API; Gemini Omni itself is built for marketing teams, video makers and developers who shape a clip one edit at a time.
