Gemini Omni 1.1 Flash: Specs, Pricing, Real Limits
Google's Gemini Omni 1.1 Flash launched today with 360p drafts, 4K output, and video editing. Verified specs, how it compares to Veo 3.1, and what a clip really costs.

Gemini Omni 1.1 Flash is Google's fast multimodal video model. It generates 10 second clips with dialogue, ambience, and effects in the same pass, accepts text, images, or an existing video as input, and spans 360p drafts to 4K output. Google shipped it on August 27, 2026 as the first major update to the Omni line announced at I/O in May (Google).
The interesting part is not the arena rank, although it holds #2 within three Elo of the leader. It is that Omni treats video as a conversation: generate a clip, then change one thing without regenerating the rest, extend it in chunks, or upload footage and edit it with a sentence. No current rival ships that loop. The catch list is real too: clips cap at 10 seconds per generation, and lip sync remains the weakest link.
Every number here comes from Google's API documentation, the live Artificial Analysis arena, published hands on reviews, and our own test generations through the API, all checked on August 27, 2026.
Gemini Omni 1.1 at a glance
The verified spec sheet, from the Gemini API documentation and the DeepMind model card:
| Spec | Gemini Omni 1.1 Flash |
|---|---|
| Maker | Google DeepMind |
| Released | August 27, 2026 (Omni line announced May 19, 2026) |
| Modes | Text to video, image to video, first and last frame, reference images, video editing, multi turn extension |
| Duration | 10 seconds per generation, extendable to 40 seconds cumulative via the API |
| Resolution | 360p draft, 720p native, 1080p and 4K via upscaling |
| Audio | Generated natively in the same pass: dialogue, ambience, effects |
| Aspect ratios | 16:9 and 9:16 |
| Open weights | No, API and Google products only, SynthID watermark on all output |
| Price (Google direct) | $0.03 to $0.30 per second depending on resolution |
What 1.1 changed over the May release: the extension context window grew from 1 to 10 seconds of prior video, cumulative length reached 40 seconds, first and last frame interpolation arrived, and the 360p draft mode landed at a third of the 720p cost and up to 60% faster. Two things Google has not published: frame rate, and any numeric benchmark scores behind its own quality claims.
One naming warning. Third party sites market this model as Veo 4. Google has never used that name: Veo remains a separate line, and Omni is positioned as Gemini reasoning fused with the rendering stack. If a spec sheet says Veo 4, treat everything else on it with suspicion.
Where it ranks right now
On the Artificial Analysis arena, where humans blind vote between clips, the picture on August 27, 2026 was:
| Ranking | Position | Elo |
|---|---|---|
| Text to video, with audio | #2, behind Wan 3.0 (1,240) | 1,237 |
| Text to video, without audio | #1 | 1,322 |
| Image to video, with audio | #4, behind MiniMax H3 Max and Seedance 2.0 | 1,179 |
| Veo 3.1, for comparison | #13 text to video | 1,089 |
Three points behind Wan 3.0 is a statistical tie, and the without audio lead suggests the picture itself votes ahead of the pack while the audio pass drags the combined score. The image to video placement matters more than it looks: Wan 3.0 does not chart there at all, so for animating stills Omni faces Seedance and MiniMax, not the arena leader.
The editing loop is the actual differentiator
Every rival generates clips. Omni also revises them. The API keeps a conversation: generate a shot, then say what changes, and the model applies the edit while holding characters, lighting, and composition steady instead of rerolling the scene. Upload an existing video and the same applies, describe the change and the rest of the footage survives. Early testers who ran dozens of generations called the physics adherence genuinely incredible on object interactions, with complex prompts landing about half the time but exceptionally when they do.
The draft economy compounds it. At $0.03 per second, a 360p draft costs a tenth of the 1080p render, so the workflow becomes: iterate the idea cheap, then rerun the winner at delivery quality. We timed it on launch day: a 10 second 360p draft came back in 25 seconds. That is fast enough to stay in the creative loop rather than context switching while you wait.
Adoption is its own signal: Adobe Firefly, Figma Weave, and Runway all shipped Omni 1.1 support on day one.
Where it falls short
The honest list, from early hands on testing across independent reviews. Lip sync holds for roughly 6 to 7 seconds in single speaker shots, then drifts, and multi speaker dialogue sync is broken outright, with lines landing on the wrong mouth. Community testing keeps landing on the same word for the motion: floaty, lacking the cinematic weight Seedance carries. Reviews of the May build found effects that feel pasted on top of scenes rather than existing inside them, and an over processed, over sharpened look.
Google's own model card concedes that consistency through edits, complex motion, and accurate text rendering remain challenges. Launch day commenters were blunter about people: uncanny valley, something robotic in the eyes. And note the regional limits: editing or extending uploaded footage is unavailable in the EEA, Switzerland, and the UK.
How to prompt Gemini Omni 1.1
Omni wants less prompt than you think. Google's own guidance is three to four sentences covering the essentials, and over specification measurably degrades output. The reliable structure, distilled from Google's own prompting guidance:
- Lead with the shot: framing plus movement in directorial vocabulary. A low angle orbit beats a low angle shot, and terms like push in, dolly zoom, and locked off all register.
- Then style, lighting, and location in one or two sentences. Keep the first prompt to four or five elements maximum.
- Describe the action plainly: who does what, in order.
- Always write at least one sound cue, since audio generates in the same pass. Quote dialogue verbatim with delivery direction, and use no music or no dialogue to suppress the defaults.
- Layer detail through follow up edits instead of one overloaded prompt. The multi turn loop is the control surface, not the first generation.
Here is that structure in practice. We ran this prompt through the API on launch day, exactly as written, and the clip below is the unedited first take. It rendered in 37 seconds at 720p.
A paper boat drifts across a puddle reflecting neon city lights at night. Gentle rain, static close shot, shallow depth of field. Soft rain patter, distant traffic hum. No dialogue, no music.
Gemini Omni 1.1 or Veo 3.1?
The question every Google user asks first, since both come from the same company. The arena answers part of it: Omni sits 148 Elo above Veo 3.1 in blind text to video voting, and a minute of 1080p costs $9 against roughly $24. Omni also brings things Veo simply does not have: the conversational editing loop, video editing from a text instruction, 360p drafts, and a 4K tier.
Veo 3.1 still earns its place in three cases. It renders 1080p natively where Omni upscales from 720p. It offers 4, 6, or 8 second durations where Omni is fixed at 10. And it can generate silent clips, while Omni's audio is always on. For an established pipeline built on exact shot lengths and native resolution, Veo holds; for everything else, Omni is the better default and the cheaper one.
What a clip actually costs
Google's direct API pricing is billed per second of output, with fixed token rates per resolution:
| Resolution | Per second | 10 second clip | Per minute |
|---|---|---|---|
| 360p draft | $0.03 | $0.30 | $1.80 |
| 720p | $0.10 | $1.00 | $6.00 |
| 1080p | $0.15 | $1.50 | $9.00 |
| 4K | $0.30 | $3.00 | $18.00 |
Hosted API platforms add a margin on top of these rates, typically 25 to 50% per second depending on the host and resolution, so the Google numbers above are the floor. For context against the field per minute of 1080p: Omni at $9, MiniMax H3 near $7.80, and Veo 3.1 around $24.
On Morphix, Gemini Omni 1.1 runs 5 credits per second at 360p, scaling to 15 at 720p, 23 at 1080p, and 45 at 4K. A 10 second draft is 50 credits, which makes the draft mode the cheapest way in the catalog to test a video idea before spending delivery money on it.
Running Gemini Omni 1.1 on Morphix
Omni 1.1 went live in the video tool on launch day, with its own model page. The integration covers text to video, image to video with start and end frames, up to 3 reference images, and video editing: attach a clip, describe the change, and Omni rewrites it. Clips are 10 seconds, 16:9 or 9:16, from 360p drafts to 4K, and audio is always on since the model has no silent mode.
Free plan credits cover a couple of 360p drafts without a card, which is exactly the mode Google built for finding out whether the model fits your work. See pricing for the full credit tiers.
Frequently asked questions
Is Gemini Omni the same as Veo 4?
No. Veo 4 is not a Google product name, just third party marketing shorthand. Omni is a separate line from Veo that pairs Gemini reasoning with Google's video rendering stack, and DeepMind still lists Veo separately for cinematic generation.
How long can Gemini Omni 1.1 videos be?
10 seconds per generation, extendable to 40 seconds cumulative through the API's multi turn interface, which reads up to 10 seconds of prior footage per extension. On Morphix each generation is a 10 second clip.
What does Gemini Omni 1.1 cost?
From Google directly: $0.03 per second at 360p, $0.10 at 720p, $0.15 at 1080p, and $0.30 at 4K, so a 10 second 1080p clip is $1.50. Hosted platforms charge a markup on those rates. On Morphix it runs 5 credits per second at 360p, up to 45 at 4K.
Does Gemini Omni 1.1 support image to video and editing?
Yes, both. It animates a still, interpolates between a first and last frame, takes up to 3 reference images for subject consistency, and edits uploaded video from a text instruction. All of these modes are live on Morphix. Note that editing uploaded footage is regionally restricted by Google in the EEA, Switzerland, and the UK.
Should I use Gemini Omni 1.1 or Veo 3.1?
Omni for most work: it ranks 148 Elo above Veo 3.1 in blind voting, costs less than half per minute of 1080p, and adds editing, drafts, and 4K. Veo 3.1 holds for pipelines that need native 1080p, exact 4, 6, or 8 second durations, or silent clips. Running the same prompt on both settles it for your footage.
Does Gemini Omni 1.1 generate audio?
Yes, always: dialogue, ambience, and sound effects generate in the same pass, and there is no silent mode. Write sound cues into the prompt to direct it, or add no music or no dialogue to suppress the defaults. Lip sync is the weak spot, holding about 7 seconds before drifting.