What Is Gemini Omni 1.1 Flash? 10-Second Video Context
Marcus Bell
Frontier Models Correspondent

TLDR10-second video context, 40-second chained extensions, first/last-frame control, 360p drafts, and up to 4K output define Gemini Omni 1.1 Flash.
Meet Gemini Omni 1.1 Flash, 10-Second Video Context
Gemini Omni 1.1 Flash is Google's preview multimodal model for generating and editing video with text, image, and video inputs. Its main update is Scene Extension, which uses up to 10 seconds of prior video context and chains 10-second additions to a cumulative 40 seconds. The model also supports first-and-last-frame control, 360p drafts, video references, sound generation, and output up to 4K. Google announced the model on August 27, 2026, with access through Google AI Studio, the Gemini API, Google Cloud's Gemini Enterprise Agent Platform, and Flow. The official launch post describes the developer-focused controls in more detail. Google's announcement
Key Takeaways
- Gemini Omni 1.1 Flash is a Google DeepMind multimodal video generation and editing model.
- Scene Extension reads up to 10 seconds of previous footage and supports chained extensions up to 40 seconds.
- First/Last-Frame Control generates motion between specified starting and ending frames.
- 360p Drafting is designed for iteration at up to 60% faster throughput and one-third of standard 720p cost.
- The model supports 360p, 720p, 1080p, and 4K video resolutions.
- Current documentation classifies the model as Preview and lists 131,072 maximum input tokens.
What Is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is the first numbered update in Google's Gemini Omni Flash family. It is designed for video creation, video editing, and multimodal creative workflows rather than general-purpose text chat. Developers can provide text, images, or video and receive video output alongside text responses.
The model extends the original Omni concept: “anything in, anything out,” beginning with video. Google describes the broader Omni family as combining Gemini's reasoning and world knowledge with generative creation. That design supports natural-language editing, reference-driven generation, and visual continuity across multiple turns.
The current model documentation gives Gemini Omni 1.1 Flash the model ID gemini-omni-1.1-flash-preview and a Preview launch stage. It lists August 27, 2026, as the release date. Google positions the update for professional production use, but Preview status means developers should still verify quality, quotas, latency, and policy behavior before depending on it for critical pipelines.
The modality details require one qualification. Google's broader Omni description includes text, image, audio, and video as native multimodal inputs. The current Gemini Enterprise Agent Platform entry for Omni 1.1 lists text, image, and video input, while marking audio input as not supported. It does list sound generation, including speech, music, and sound effects, as supported.
Gemini Omni 1.1 Flash at a Glance
| Specification | Current information |
|---|---|
| Developer | Google DeepMind |
| Type | Multimodal generative video generation and editing model |
| Model ID | gemini-omni-1.1-flash-preview in current Agent Platform documentation |
| Input modality | Text, image, and video; audio input is marked not supported in the current Agent Platform entry |
| Output modality | Video and text; sound generation is supported |
| Context window | 131,072 maximum input tokens; 57,920 maximum output tokens |
| Video context for extension | Up to 10 seconds of preceding video |
| Extension length | 10-second increments, up to 40 seconds cumulative |
| Resolutions | 360p, 720p, 1080p, and 4K |
| Video references | Up to 3 seconds of reference video; current documentation allows up to 3 videos per prompt |
| Pricing | $1.50 / 1M input tokens; $9 / 1M text-output tokens; $0.10 / second of 720p video output |
| Availability | Preview access through Google AI Studio, Gemini API, Agent Platform, Flow, and kie.ai API |
| License | Not yet confirmed |
| Open weights | Not yet confirmed |
The published cloud pricing table identifies the video charge as $0.10 per second, equivalent to $17.50 per 1 million video-output tokens under its stated tokenization. Google Cloud's pricing documentation
How Gemini Omni 1.1 Flash Works and What Makes It Different
The update is best understood as a control layer for generative video. A conventional text-to-video request produces a clip from a prompt. Omni 1.1 adds ways to constrain how a shot starts, how it ends, what it should remember, and which visual references it should preserve.
Gemini Omni 1.1 Flash extends scenes by reading up to 10 seconds of prior video context and chaining 10-second additions to 40 seconds total.
Scene Extension
Scene Extension continues an existing clip from its endpoint. Omni 1.1 can analyze up to 10 seconds of preceding footage, compared with the final second in the earlier comparison described by Google's launch materials. This broader context can help preserve a character's appearance, camera direction, environment, lighting, and narrative state.
Extensions happen in 10-second increments and can reach 40 seconds cumulatively. That figure describes a sequence of linked generations, not a single 40-second render. The model appends footage to the end of a clip. It does not support prepending or inserting new footage into the middle of an existing clip through this feature.
Stateful editing is another part of the workflow. Through the Interactions API, a developer can refer to a previous interaction and describe the next change. The model then applies the instruction while preserving details the prompt does not ask it to change. This reduces the need to upload the entire prior video for every conversational edit.
First/Last-Frame Control
First/Last-Frame Control lets a developer provide the opening and closing frames of a shot. Omni 1.1 generates the continuous motion between those two keyframes.
This is useful for controlled camera movement, including an orbit around a subject, a zoom between compositions, or a loop designed to return to its starting visual state. The capability shifts part of the creative process from describing an outcome to specifying boundary conditions.
Video References
Video Reference Input lets a prompt include short footage as a guide for movement, appearance, or consistency. Google's announcement describes references of up to 3 seconds. Current documentation allows up to 3 videos per prompt, although the provided technical material warns that reasoning across multiple videos may degrade output.
A reference clip is not necessarily copied frame by frame. It supplies additional visual context for the generated result. Audio inside a video reference is ignored according to the available technical release documentation.
Draft-to-4K Workflow
The Draft-to-4K Workflow separates rapid experimentation from final rendering. Google says 360p drafts generate up to 60% faster than 720p output and cost one-third as much as standard 720p generation. Developers can test prompts and shot direction at lower resolution, then select a result for 1080p or 4K output.
The bundle does not confirm whether every upscale preserves the exact same motion, random seed, hand details, or facial details from a 360p draft. That question matters when a team evaluates a draft-then-final pipeline.
What You Can Do With Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash is suited to workflows where a short generated clip must remain editable. Practical uses include:
- Extend narrative scenes: Continue a conversation, camera move, or action sequence while retaining up to 10 seconds of prior visual context.
- Storyboard controlled shots: Give the model a first frame and last frame, then generate the transition between them.
- Create seamless loops: Use matching boundary frames for product loops, social clips, interface backgrounds, or visual installations.
- Prototype creative directions: Generate 360p drafts quickly, compare multiple prompts, and reserve higher-resolution output for selected concepts.
- Maintain character references: Add short video references to guide likeness, movement, or visual style.
- Build conversational editors: Use natural-language instructions to change an environment, camera angle, object, or action across multiple turns.
- Produce explainers: Combine Gemini's world knowledge with generated visuals to make complex scientific, historical, or cultural ideas easier to show.
- Support media-production tools: Integrate video generation and editing into creative applications through the Gemini API or an enterprise agent workflow.
These uses reflect Google's launch examples and the documented capabilities. They do not guarantee reliable character continuity or prompt adherence in every scene.
How Gemini Omni 1.1 Flash Compares
The clearest comparison in the bundle is with Veo's scene-extension context. It concerns continuation context, not overall video quality, cost, or benchmark performance.
| Capability | Gemini Omni 1.1 Flash | Veo comparison in Google's launch material |
|---|---|---|
| Prior video context for scene extension | Up to 10 seconds | 1 second |
| Extension approach | 10-second increments, up to 40 seconds cumulative | Not yet confirmed |
| First-and-last-frame control | Supported | Not yet confirmed |
| Video references | Supported, up to 3 seconds described by Google | Not yet confirmed |
For scene extension, Google's launch comparison gives Omni 1.1 a 10-second context window versus 1 second for Veo. This is a narrow capability comparison and should not be read as a general ranking between the models. Google's comparison post
Availability: How to Access Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash is a hosted Preview model. The available evidence does not show downloadable weights or a self-hosting package.
| Access route | What is confirmed |
|---|---|
| Google AI Studio and Gemini API | Google announced developer access for generative video workflows and conversational editing. |
| Gemini Enterprise Agent Platform | Google Cloud documentation lists the Preview model, global availability, fixed quota support, and Agent Studio access. |
| Google Flow | Google DeepMind announced a Flow rollout for trying the model and related creative controls. |
| kie.ai API | kie.ai offers Gemini Omni 1.1 Flash via API, giving developers a hosted integration route for the model. |
| Downloadable weights | Not yet confirmed |
| Self-hosting | Not yet confirmed |
The current Agent Platform documentation marks pay-as-you-go as not supported and fixed quota as supported. This may reflect that product surface's Preview configuration rather than every Gemini API billing path. Developers should check the applicable console and account terms before estimating production capacity.
Google's official documentation names the Interactions API as the route for stateful conversational workflows. The available model entry also lists structured output, function calling, grounding, code execution, and tuning as unsupported. That makes Omni 1.1 primarily a media-generation component, not a general tool-using agent model.
What We Don't Know Yet
Several important details remain open:
- Independent benchmark validity: Community posts claimed a score of 1,515 and first place in a text-to-video Arena, while other posts reported 1,488 or fourth place with a score of 76.24. The posts do not establish a shared leaderboard, evaluation method, sample size, or comparable conditions. These claims are unconfirmed. One reported Arena result
- Final pricing by resolution: Google confirms that 360p costs one-third as much as standard 720p and publishes $0.10 per second for 720p video output. Exact 360p, 1080p, and 4K billing behavior is not fully confirmed in the supplied pricing material.
- Flow quotas: A community post claimed that Gemini Pro subscribers could generate 100 ten-second videos per month in Flow. That is a single-source report and is unconfirmed.
- Long-scene reliability: The 40-second cumulative limit is documented, but consistent quality across four linked extension stages has not been independently established.
- Prompt adherence: Early community testing was positive about clarity and music, but one tester reported difficulty getting the intended scene after roughly five prompt variants. Results remain workload-dependent.
- Audio input behavior: The broader Omni description includes audio among its modalities, while the current 1.1 Agent Platform entry marks audio input unsupported. The exact availability of audio input may vary by product surface.
- General availability and licensing: The model remains documented as Preview. Open weights, an open-source license, and a final general-availability date are not confirmed.
The available evidence supports a controllability update, but it does not yet prove consistent 40-second production scenes across prompts.
Frequently Asked Questions
What is Gemini Omni 1.1 Flash?
Gemini Omni 1.1 Flash is Google's preview multimodal model for generating and editing video with text, image, and video inputs. Its defining update is scene extension using up to 10 seconds of prior video context, alongside first/last-frame control, 360p drafting, video references, and output up to 4K.
Is Gemini Omni 1.1 Flash open source?
Gemini Omni 1.1 Flash is not confirmed as open source. The available documentation describes hosted access through Google's developer products and does not publish model weights or an open-source license.
How much does Gemini Omni 1.1 Flash cost?
Published Gemini Omni pricing lists $1.50 per 1 million input tokens, $9 per 1 million text-output tokens, and $0.10 per second of 720p video output. Google describes 360p drafts as costing one-third as much as standard 720p output, while exact 1.1 billing treatment for every resolution is not fully confirmed.
How long can Gemini Omni 1.1 Flash videos be?
Gemini Omni 1.1 Flash can extend videos in 10-second increments to a cumulative length of 40 seconds. The model uses up to 10 seconds of preceding video context for scene continuation, rather than generating 40 seconds in one render.
Where can I use Gemini Omni 1.1 Flash?
Developers can use Gemini Omni 1.1 Flash through Google AI Studio and the Gemini API, while Google Cloud documents it on the Gemini Enterprise Agent Platform. Flow is another official route, and kie.ai offers Gemini Omni 1.1 Flash through its API service.
Gemini Omni 1.1 Flash vs Veo: what is the difference?
Gemini Omni 1.1 Flash is positioned around controllable generation and editing, including first/last-frame inputs and conversational scene extension. Google's launch comparison says Omni 1.1 can use up to 10 seconds of video context for extension, compared with 1 second for Veo; this does not establish a general quality ranking.
What to watch next is Google's transition from Preview to general availability, clearer resolution-specific pricing and quotas, and independent tests of four-stage scene continuity and prompt adherence.
Building similar controllable AI video workflows? On kie.ai you can try Gemini Omni 1.1 Flash, Gemini Omni, and Wan 3.0 Video.
About Marcus Bell
Marcus reports on frontier model launches and leaks, weighing community testing against official specs.
View all posts by Marcus Bell