What Is LTX-2.5? Open-Weights Video Model

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: August 13, 2026
LTX-2.5 open-weights video model by Lightricks

TLDR10s video in 6.8s on two GB200s, 22B open weights, free under $10M revenue. LTX-2.5 is Lightricks' video world model.

LTX-2.5 Is Lightricks' Open-Weights World Model — 10s Video in 6.8s

LTX-2.5 is an open-weights video generation model from Lightricks that turns text, image, and video inputs into synchronized video and audio in a single model. It released on August 11, 2026 as a 22B-parameter diffusion model, and Lightricks reports it generates a 10-second image-to-video clip in 6.8 seconds on two NVIDIA GB200 superchips, or 23.7 seconds through the LTX API. Commercial use is free for companies under $10M in annual revenue under the LTX-2.x Community License.

Updated 2026-08-17: independent testers confirmed LTX-2.5 runs fast on consumer and Apple Silicon hardware, plus a multi-subject reference LoRA and broad tool integrations landed within a week (see the Update below).

Key Takeaways

  • LTX-2.5 is an open-weights video-and-audio model from Lightricks, released August 11, 2026 with 22B distilled and full checkpoints on Hugging Face.
  • It accepts text, image, and video inputs and generates picture plus synchronized audio in one model, no separate audio stage.
  • Published speed: a 10-second clip renders in 6.8 seconds on two GB200s and 23.7 seconds via API, versus 52–70 seconds for the fastest closed rivals.
  • Native multishot generates connected shots in one pass, holding character, scene, lighting, and voice consistent across cuts.
  • Free commercial use under a $10M annual-revenue threshold; above that requires a paid license.
  • API pricing starts at $0.09/second at 720p according to third-party posts, unconfirmed in the official launch post.

What Is LTX-2.5?

LTX-2.5 is a video foundation model built by Lightricks, the company behind the LTX video family and LTX Studio. It generates high-fidelity video with synchronized audio from three input types: text, image, and video. Dialogue, music, and ambient sound are produced together with the visuals inside the same model.

The model is described on its Hugging Face card as "an open world model with open weights, built for local execution and fine-tuning." Its primary use is generating synchronized video and audio; the card notes that applicability to "emerging domains such as robotics and physical AI is developing."

LTX-2.5 is a released, downloadable model, not a leak or preview. Weights, inference code, and training code went public at launch. It comes in two API variants: ltx-2-5-fast for speed and low cost, and ltx-2-5-pro for higher fidelity. Fast produces portrait and landscape video up to 4K; Pro tops out at 1080p.

LTX-2.5 at a Glance

AttributeDetail
DeveloperLightricks
TypeOpen-weights video generation / world model
Parameters22B (distilled and full checkpoints)
ModalityText-to-video, image-to-video, video-to-video, audio-to-video, all with synchronized audio
Max resolutionUp to 4K (Fast); 1080p (Pro)
Duration6–20s at 720p/1080p (24/25fps); 6–10s at 1440p/4K
Published speed6.8s for a 10s clip on 2× GB200; 23.7s via API
Min VRAM16GB (reported by third party, unconfirmed)
API pricingFrom $0.09/s (720p) per third-party posts; unconfirmed
LicenseLTX-2.x Community License — free under $10M revenue
AvailabilityOpen weights on Hugging Face; LTX API; LTX Studio

How LTX-2.5 Works / What Makes It Different

LTX-2.5 is described by a third-party summary as a 22B-parameter asymmetric dual-stream diffusion transformer. According to Lightricks researcher Yoav Hacohen, the model uses Variable Token Rate (VTR) and pixel diffusion, with detailed results reported on a single GPU (@yoavhacohen, Aug 11, 2026). Several architecture pieces stand out.

Diffusion Fidelity Rendering allocates rendering compute by scene complexity instead of locking every scene to one compression rate. Lightricks describes it as delivering "industry-leading pixel quality that holds up frame by frame, even on a cinema screen." The model renders scenes from a grid of high-fidelity keyframes.

Native multishot is the headline change from earlier versions. A single generation produces multiple connected shots that hold character identity, environment, lighting, voice, and visual style across cuts. Previous versions produced a single continuous shot.

A new Diffusion Video Decoder replaces the older VAE reconstruction stage, which the model card credits with sharper faces, textures, and on-screen text. A custom Gemma 4 12B text encoder holds complex prompts together across longer sequences, and an optional Duration Predictor node predicts a clip's length from the prompt and sets the frame count automatically.

"LTX-2.5 predicts the next moment, not the next word" is how one launch summary framed the world-model angle. That framing captures why Lightricks positions it beyond a straight text-to-video tool.

What You Can Do With LTX-2.5

The API exposes several endpoints. Text-to-video takes a scene description and returns a complete clip with matching audio, up to 4K and 20 seconds per request. Image-to-video animates a still with motion, depth, and audio while preserving the source's visual identity. Audio-to-video generates visuals synchronized to a supplied dialogue, music, or ambient track.

Additional workflow endpoints include Retake (regenerate a specific section without starting over), Extend (lengthen a clip from either end with seamless continuity), and Reframe (change a video's resolution with generated fill). These editing endpoints run on the LTX-2.3 tier per the current support matrix, not the 2.5 variants.

Because weights are open, teams can self-host and fine-tune the raw foundation on their own data. Day-zero support landed for Diffusers and ComfyUI, and community projects including Ostris AI Toolkit and WanGP added support within a day of release. If you want to prototype a comparable image-to-video pipeline through a hosted API instead of self-hosting, models like MiniMax H3 cover a similar text-and-image-to-video task surface.

How LTX-2.5 Compares

Speed is where LTX-2.5's published numbers separate from rivals. The comparison below uses Lightricks' own measured figures for a 10-second image-to-video clip; independent tests vary widely by hardware and configuration.

Model10s clip generationWeights
LTX-2.5 (on-prem, 2× GB200)6.8sOpen
LTX-2.5 (API)23.7sOpen
Veo 3.170s (8s clip)Closed API
MiniMax H3180sOpen, conditional
Seedance 2.5317sClosed API
Kling 3.0 Pro398sClosed API

Quality leadership is disputed. In a single-user RTX 5090 comparison, one tester clocked LTX-2.5 at roughly 1.5 minutes versus 8–10 minutes for MiniMax H3, while rating H3 better for quality, complex motion, consistency, and native audio-visual output (@FiniYang, Aug 13, 2026). Other testers echoed that speed-versus-quality split, and one community member alleged the official comparison table was misleading. For a closer look at a competing single-shot approach, see our analysis of what Seedance 2.5 is.

Availability: How to Access LTX-2.5

LTX-2.5 open weights are published on Hugging Face at the Lightricks/LTX-2.5 model card, gated behind a contact-information agreement. Inference and training code are available through Lightricks' official GitHub and documentation.

The hosted LTX API serves both ltx-2-5-fast and ltx-2-5-pro variants through sync and async endpoints, documented in the official LTX docs. LTX-2.5 is also available on all LTX Studio plans, including the free tier with a one-time 800-credit allocation.

Third-party API prices circulating after launch list $0.09/second at 720p, $0.15 at 1080p, $0.19 at 2K, and $0.37 at 4K (@GenAI_Creative, Aug 12, 2026). These are not stated in the official launch post and remain unconfirmed against official pricing.

What We Don't Know Yet

Several claims remain single-source or unresolved as of August 13, 2026:

  • Benchmark comparability. The 6.8-second figure uses two GB200s. Independent results range from about 22 seconds on an RTX Pro 6000 to roughly 7 minutes of generation on an A100, with no common resolution, frame count, or checkpoint across tests.
  • Commercial terms. The under-$10M revenue threshold and per-second API prices come from third-party posts; the official launch post does not state them.
  • Capability specifics. Native audio up to 1080p, 4K HDR output, LoRA/IC-LoRA compatibility, and 16GB operation are each reported by individual accounts without corroborating tests in the supplied posts.
  • Local setup friction. One RTX 4090 test found a Torch/CUDA mismatch degrading throughput before a fourfold correction; an A100 Colab test needed roughly 2.5 hours of setup before a ~7-minute generation run.

Frequently Asked Questions

What is LTX-2.5?

LTX-2.5 is an open-weights video generation model from Lightricks that produces synchronized video and audio from text, image, and video inputs. It released on August 11, 2026 with 22B distilled and full checkpoints and runs locally or through the LTX API.

Is LTX-2.5 open source?

LTX-2.5 ships with open weights under the LTX-2.x Community License. Commercial and production use is free for companies with under $10M in annual revenue; above that threshold a paid license is required, and transfer of fine-tunes may also need a paid license.

How fast is LTX-2.5?

LTX-2.5 generates a 10-second image-to-video clip in 6.8 seconds on two NVIDIA GB200 superchips on-premises, and in 23.7 seconds through the LTX API, according to Lightricks' published figures. Independent hardware tests report slower, configuration-dependent times.

How much does the LTX-2.5 API cost?

According to third-party pricing posts, LTX-2.5 API pricing starts at $0.09 per second at 720p, $0.15 at 1080p, $0.19 at 2K, and $0.37 at 4K. These figures are not stated in the official launch post and remain unconfirmed.

What hardware do you need to run LTX-2.5?

LTX-2.5 is built for local execution, and a third-party summary reports a 16GB minimum VRAM requirement. Community testers ran it on RTX 4090, RTX 5090, and A100 GPUs, though performance depended heavily on Torch/CUDA configuration.

What is native multishot in LTX-2.5?

Native multishot lets LTX-2.5 generate multiple connected shots in a single pass, holding character, environment, lighting, voice, and visual style consistent across cuts. Previous LTX versions produced a single continuous shot instead.

How does LTX-2.5 compare to MiniMax H3?

LTX-2.5 is substantially faster than MiniMax H3 in community tests, with one RTX 5090 run showing 1.5 minutes versus 8–10 minutes. Some testers rate H3 higher for complex motion and overall quality, and Lightricks' own comparison table has been disputed by some community members.

What to Watch Next

Three signals will clarify LTX-2.5's real standing. First, watch for a standardized independent quality benchmark against MiniMax H3 and Seedance 2.5 using a common resolution and checkpoint, since current comparisons disagree. Second, track whether Lightricks confirms the third-party API prices and the $10M license threshold in official documentation. Third, follow community fine-tuning and consumer-GPU results, where 16GB operation and RTX-class performance still look configuration-dependent.

Update — 2026-08-17

In the days after launch, independent testers filled in some of the consumer-hardware gaps the launch numbers left open. On an RTX 5090, one hands-on report generated a 10-second Full HD image-to-video clip in about two minutes (@hectorVFX), while a local 480×640 comparison clocked LTX-2.5 at 45 seconds using an eight-step distilled model, against 63 seconds for MiniMax H3 and two to three minutes for cloud Seedance 2.5 — though the tester cautioned the timings were not strictly controlled (@yu_ichi_suzuki). Apple Silicon results also emerged: on an M5 Max MacBook Pro with 128GB RAM via MLX, five-second generation ranged from roughly 30 seconds (Q4) to 45 seconds (BF16), with peak Metal RAM between 16.3GB and 42.0GB depending on precision (@trevorwood222). These figures give a rough precision-versus-memory picture the official launch numbers omitted, but still fall short of a consistent VRAM/RAM-by-resolution matrix.

The quality-versus-speed split the article flagged persisted in new head-to-heads. A same-seed crash-scene comparison put LTX-2.5 at roughly 60–90 seconds versus about six minutes for MiniMax H3, but rated H3 higher for interpreting impacts (@BennyDaBall_OG), consistent with earlier reports that LTX-2.5 wins on speed while rivals hold an edge on complex motion.

Tooling and community integrations widened quickly. Within the release week, community posts reported support across ComfyUI, Ostris AI Toolkit, Maestro, WanGP, GGUF, MLX, fal, and Chutes, and a multi-subject reference LoRA appeared that accepts up to five reference images while preserving multiple characters, clothing, objects, and backgrounds (@SD_Tutorial). None of this resolves the still-open questions around an official benchmark table, license terms, or the single-sourced native-audio and 1080p claims.

Building similar text-to-video or image-to-video pipelines? On kie.ai you can try MiniMax H3, Seedance 2.5, and Wan3.0-Video.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair