Grok Imagine Video 1.5: Release Deep Dive
Marcus Bell
Frontier Models Correspondent

TLDRGrok Imagine Video 1.5 launched in the API on June 16, with a July 31 update adding native 1080p, text-to-video, seven visual references, voice references, and native synchronized audio.
Grok Imagine Video 1.5: Deep Dive Into the 1080p, Multi-Reference, and Native-Audio Update
On August 1, AI analyst Rohan Paul summed up the latest change: "text-to-video support, image and voice references, and native 1080p arrived on Grok Imagine Video 1.5" (@rohanpaul_ai). That post described a July 31 rollout on a model that was no longer merely preview software: xAI had announced Grok Imagine Video 1.5 as generally available in the API on June 16, 2026. The launch and subsequent update now give builders a clearer picture of what is available, what is documented, and where the numbers remain incomplete.
TLDR Grok Imagine Video 1.5 is xAI's launched image-to-video and text-to-video model. Its July 31 update added native 1080p, image and voice references, and text-to-video support, with up to seven simultaneous visual anchors per generation. It also generates synchronized sound effects, ambience, and dialogue in the same pass. The xAI API lists output pricing of $0.08 per second at 480p, $0.14 at 720p, and $0.25 at 1080p. An Image-to-Video Arena snapshot records 1460 points, ahead of FLUX 3 Video and behind Gemini Omni Flash.
Key Takeaways
- Grok Imagine Video 1.5 launched as a generally available model in the xAI API on June 16, 2026. Its model name is
grok-imagine-video-1.5. - A July 31 update added text-to-video, native 1080p, image references, voice references, and multi-reference scene control, with up to seven simultaneous visual anchors per generation.
- xAI's official announcement says sound effects, ambience, and dialogue are generated in the same pass as video, with speech that is clearer and better synced.
- xAI documentation lists resolution-tiered API pricing per second: $0.08 at 480p, $0.14 at 720p, and $0.25 at 1080p, plus $0.01 per input image.
- The current Image-to-Video Arena evidence records 1460 points for Grok Imagine Video 1.5, placing it behind Gemini Omni Flash at 1462 and ahead of FLUX 3 Video at 1453. Older third-party posts that called it #1 should not be treated as the current ranking.
- The exact duration and resolution combinations for every voice-reference and multi-reference workflow remain incompletely specified.
What Actually Shipped on August 1
The July 31 rollout, described publicly on August 1, was a capability expansion rather than a new model version. Grok Imagine Video 1.5 had already reached general availability in the API on June 16. The later update widened what the model can do across the API, web, and mobile workflows.
Four additions dominate the current signal.
First, native 1080p. Text-to-Video and Image-to-Video now support native 1080p, as described by convergent posts from @fal, @AskVenice, and @teslaownersSV. The broader API documentation also lists 480p, 720p, and 1080p pricing tiers. What is not fully pinned down is whether every combination of reference inputs exposes the same resolution choices.
Second, Reference-to-Video and voice references. The updated workflow can use image references to steer a scene and voice references to support face and voice consistency. The strongest confirmed ceiling in the current signal is up to seven simultaneous visual anchors or image references per generation. The available sources do not establish a separate maximum number for voice references, so the old claim of a fixed audio-reference count should not be treated as confirmed.
Third, Text-to-Video support was added to the 1.5 generation in this update, per @fal. Earlier, the model was primarily described as an image-first tool. The addition means you can now start from a prompt alone, not just a still frame.
Fourth, the update broadened access to these capabilities. Text-to-video and native 1080p were made generally available on grok.com/imagine, iOS, and Android. Image and voice references started on SuperGrok Heavy and SuperGrok Plus through web and iOS before rolling out more broadly. The model is also available through the xAI API and third-party developer infrastructure, including Vercel AI Gateway and WaveSpeed.
The through-line across all three generation modes is Single-Pass Audio: synchronized sound effects, ambience, and dialogue generated in the same inference step as the video. xAI's own announcement describes speech that is "clearer and better synced" and lands "on the action" (x.ai news post).
The launch also included a speed-focused 1.5 Fast variant for Grok's web and mobile experience. xAI says it produces 6-second, 720p videos in about 25 seconds, down from more than 40 seconds with the previous model (x.ai news post).
The Coined Terms Worth Tracking
A few capabilities in this release have crisp enough definitions to be worth naming precisely, because they are what downstream coverage will reference:
- Reference-to-Video (R2V) is the workflow that uses image references to steer character, style, and composition across a generated clip. The updated system supports up to seven simultaneous visual anchors.
- Single-Pass Audio is xAI's approach of generating video and synchronized audio in one inference step, with no separate audio stage or post-production pass required for the generated sound.
- Multi-Reference Control is the use of several image anchors in one generation to guide a scene, subject, or visual identity. Seven simultaneous visual anchors is the supported figure in the current update signal.
- Voice References are inputs intended to help preserve face and voice consistency. The feature is available in the updated workflow, but the sources here do not establish a separate maximum number of voice references.
- 1.5 Fast is the speed-tuned variant xAI cites as producing 6-second, 720p videos in about 25 seconds, down from 40-plus seconds on the previous model (x.ai news post).
Single-Pass Audio remains the clearest differentiator here: it folds sound design into the generation step, so a clip can arrive with effects, ambience, and dialogue already aligned to the motion. That capability belongs to the launched model itself; it was not newly introduced only by the July 31 update.
The Pricing Picture
Pricing is one of the few areas with a primary vendor number. xAI's model documentation lists per-second output pricing scaled by resolution: $0.08 per second at 480p, $0.14 at 720p, and $0.25 at 1080p, with an additional $0.01 per input image (docs.x.ai). Audio input using preset voices is listed as free.
That documentation also pins down identity details worth recording: the model name is grok-imagine-video-1.5, with aliases grok-imagine-video-1.5-preview and grok-imagine-video-1.5-2026-05-30, served from the us-east-1 and us-west-2 regions at a rate limit of 10 requests per second (docs.x.ai).
The previously circulating flat $0.080-per-second figure is not a replacement for the documented tiered price. It corresponds to the 480p output tier in xAI's documentation. For API budgeting, use the resolution-specific prices rather than treating $0.080 as a universal rate.
If you want to run a controlled cost-per-clip comparison against a similarly positioned image-to-video model before committing, Seedance 2.5 exposes the same image-to-video task type and is a reasonable baseline for that kind of eval.
Grok Imagine Video 1.5 vs Seedance 2.0: What the Signal Says
The comparison that recurs across the signal set is against ByteDance's Seedance 2.0. Earlier third-party pages and creators framed Grok Imagine Video 1.5 as having taken the top spot on a crowd-sourced Image-to-Video Arena. The newer Arena evidence gives a more precise snapshot: Grok Imagine Video 1.5 scored 1460 points, just behind Gemini Omni Flash at 1462 and ahead of FLUX 3 Video at 1453. The available sources do not establish one universal ranking across every historical post or leaderboard view.
Here is how the comparison holds up dimension by dimension.
- Arena ranking. An Arena snapshot records Grok Imagine Video 1.5 at 1460 points, behind Gemini Omni Flash and ahead of FLUX 3 Video. Older third-party writeups claimed that Grok reached #1, and a hands-on reviewer echoed the framing that 1.5 "took the #1 spot… knocking Seedance 2.0 out" (a YouTube review). Those older claims are not the best description of the current recorded position.
- Max resolution. Grok Imagine Video 1.5 supports native 1080p after the July 31 update. The available signal does not provide a comparable Seedance 2.0 resolution number: unverified — no public number from this signal set.
- Native audio. Grok Imagine Video 1.5 generates synchronized audio in the same pass, per xAI. The signal set carries no equivalent Seedance 2.0 audio specification: unverified — no public number from this signal set.
- Community verdict. Arena positions and hands-on impressions do not always align. On a Reddit thread, one user said Grok was "not even as good as Seedance even though they claimed to be better," and another agreed Seedance "has a je ne sais quoi" (r/grok discussion).
The honest read: Grok Imagine Video 1.5 now has a recorded 1460-point Arena result, but the bundle does not support describing it as the uncontested leader. The score gives a useful current data point; the divergent hands-on opinions are still a reason to run your own side-by-side before routing production traffic.
What We Know vs. What We Don't
Separating the confirmed from the circulating matters more than usual here, because the launch and the later feature rollout were documented across official pages, platform documentation, and third-party accounts.
What we know:
- xAI announced Grok Imagine Video 1.5 as generally available in the API on June 16, 2026. The API model name is
grok-imagine-video-1.5(x.ai). - The July 31 update added text-to-video, native 1080p, image references, voice references, and multi-reference scene control. The update was described publicly on August 1 (@rohanpaul_ai).
- Up to seven simultaneous visual anchors or image references are supported per generation. Voice references are also supported, although a separate maximum for them is not established in the available sources.
- Text-to-video and native 1080p are generally available on grok.com/imagine, iOS, and Android. Image and voice references began rolling out through web and iOS plans, with broader availability described in the update signal.
- Grok Imagine Video 1.5 generates synchronized sound effects, ambience, and dialogue in the same inference pass, per xAI's official announcement (x.ai).
- xAI documentation lists per-second pricing of $0.08 at 480p, $0.14 at 720p, and $0.25 at 1080p, plus $0.01 per input image (docs.x.ai).
- The model is served from
us-east-1andus-west-2at 10 requests per second (docs.x.ai). - The Image-to-Video Arena has a recorded score of 1460 points, behind Gemini Omni Flash at 1462 and ahead of FLUX 3 Video at 1453.
What we don't know:
- Whether every combination of text, image, voice, and multi-reference inputs supports the full 1080p option. Native 1080p is confirmed for the update, but the available sources do not provide a complete mode-by-mode resolution matrix.
- Whether every updated reference workflow has the same duration range. Available API documentation describes 1-to-15-second generation generally, while exact limits for each reference combination remain unclear.
- The precise plan-by-plan rollout status for image and voice references. The update began with SuperGrok Heavy and SuperGrok Plus on web and iOS before broader rollout, but the bundle does not provide a complete current availability matrix.
- Whether the older #1 Arena claims refer to a different leaderboard snapshot or ranking configuration. The current evidence records 1460 points and a position behind Gemini Omni Flash, while other accounts describe Grok as #2 or in the top three.
- Whether an official benchmark suite will be published. The available sources identify the crowd-sourced Arena result but not a separate official benchmark suite.
- How consistently the new reference controls preserve identity and voice across different prompts, subjects, and production conditions. Community quality assessments remain mixed, with positive hands-on claims contrasting with reports of unusable outputs in some tests.
Why This Matters for Builders
For anyone routing video generation in production, the interesting shift is architectural, not cosmetic. Single-Pass Audio removes a stage. A pipeline that previously generated silent clips and then bolted on a separate audio model can, in principle, collapse into one call. That changes both latency budgeting and the failure surface, because audio and motion now succeed or fail together.
Multi-Reference Control changes a different workflow. Teams that manually chained frames to hold a character's face across shots now have up to seven visual anchors for that consistency. Whether it works well enough to replace manual continuity work is exactly the kind of claim you should test on your own assets rather than take from an Arena score.
Native 1080p changes the delivery calculation. The earlier article's simple 720p-versus-1080p split for reference mode is no longer reliable: the July 31 update added native 1080p, but the available sources do not fully specify which resolution combinations are exposed when multiple image or voice references are active. That leaves a practical routing question rather than a settled hard cap. Builders should verify the exact mode, plan, and resolution combination their workflow needs.
The API is no longer a preview-only integration target. The launched model uses the stable grok-imagine-video-1.5 name, and the same model is exposed through third-party infrastructure such as Vercel AI Gateway and WaveSpeed. That makes it possible to compare direct xAI API costs with gateway costs while keeping the generation behavior under test.
How to Evaluate It Yourself
The gap between the Arena result and the hands-on community dissent is the reason to run your own test rather than trust either number. A few concrete checks:
Pin the resolution question first. Generate the same prompt in text-to-video and image-to-video at 1080p, then repeat with your intended reference configuration. Confirm which resolution choices are actually exposed for your account, API route, and reference combination. Do not assume that a general 1080p announcement answers every multi-reference edge case.
Stress the audio. xAI says that sound effects, ambience, and dialogue are generated in the same pass and that speech is better synced. Test dialogue, action-linked effects, and multi-character voice consistency directly, because you cannot swap out the audio stage if it underperforms.
Measure the cost per usable clip. Use the documented $0.08, $0.14, and $0.25 per-second API tiers, add the $0.01 input-image charge where applicable, and record retries separately. The flat $0.080 figure should not be used as a universal price.
Probe the action and continuity ceiling. Community quality assessments remain disputed, especially in comparisons with Seedance. Test fast movement, faces, scene transitions, and repeated generation from the same reference set. The seven-anchor limit describes how many visual controls can be supplied; it does not guarantee perfect continuity.
Finally, test availability rather than relying on a plan label. Text-to-video and native 1080p are generally available on grok.com/imagine, iOS, and Android, while image and voice references have a staged rollout history. Confirm the features shown in the specific web, mobile, or API environment you plan to ship.
What to Watch Next
Most of the original open questions have now been answered. The launch date, API status, model name, core update features, native 1080p addition, seven-visual-reference ceiling, and tiered API pricing are documented well enough to use in a current model comparison.
The remaining signals are narrower. Watch for a complete mode-by-mode resolution and duration matrix, especially for multi-reference and voice-reference requests. Watch for a definitive current Arena leaderboard view that reconciles the older #1 language with the recorded 1460-point result. And watch for an official benchmark suite or model card that goes beyond crowd-sourced rankings.
For builders, the most consequential next test is not whether Grok Imagine Video 1.5 exists or whether it is still in preview; it is how reliably the launched model combines 1080p delivery, seven visual anchors, voice references, and synchronized audio in the same production workflow.
Building similar image-to-video, video-editing, and native-audio workflows? On kie.ai you can try Seedance 2.5, Kling O3, and Gemini Omni 1.1 Flash.
Frequently Asked Questions
When did Grok Imagine Video 1.5 launch, and where is it available?
xAI announced Grok Imagine Video 1.5 as generally available in the API on June 16, 2026. The July 31 update made text-to-video and native 1080p generally available on grok.com/imagine, iOS, and Android, while image and voice references began rolling out through web and iOS plans.
Does Grok Imagine Video 1.5 support native 1080p output?
Yes. The July 31 update added native 1080p for text-to-video and image-to-video. The exact resolution combinations available in every reference mode are not fully specified in the available documentation.
How many reference images does Grok Imagine Video 1.5 accept?
The updated workflow supports up to seven simultaneous visual anchors or image references per generation. Voice references are also supported, but the available sources do not establish a separate maximum number for them.
Does Grok Imagine Video 1.5 generate audio?
Yes. Grok Imagine Video 1.5 generates sound effects, ambience, and dialogue in the same pass as the video. xAI says speech is clearer and better synced and that audio lands on the action.
How much does Grok Imagine Video 1.5 cost through the xAI API?
xAI's documentation lists output pricing of $0.08 per second at 480p, $0.14 per second at 720p, and $0.25 per second at 1080p. Image input is listed at $0.01 per image, while audio input using preset voices is listed as free.
What is the maximum clip length?
Available API documentation describes a configurable duration from 1 to 15 seconds. The exact duration limits for every combination of text, image, voice, and multi-reference inputs are not fully specified.
Have official benchmarks been published for Grok Imagine Video 1.5?
No official benchmark suite is identified in the available sources. The Image-to-Video Arena has a recorded score of 1460 points, placing Grok Imagine Video 1.5 behind Gemini Omni Flash at 1462 and ahead of FLUX 3 Video at 1453.
About Marcus Bell
Marcus reports on frontier model launches and leaks, weighing community testing against official specs.
View all posts by Marcus Bell