LTX 2.5: multi-shot scenes with sound baked in, from an open-weights model
Updated on August 16, 2026
LTX 2.5 is an open-weights video generation model from LTX, the company spun out of Lightricks. Released on August 11, 2026, it renders picture and synchronized sound in a single pass, working from text, a still image or an existing clip. The weights cost nothing on Hugging Face, while the hosted API bills every generated second. Filmmakers, ad teams and developers with a capable GPU under the desk get the most out of it.
- Picture and audio rendered together, lip-sync included
- Native multi-shot with stable characters across cuts
- Open weights, runs locally from 16 GB of VRAM
- Faster-than-real-time rendering on high-end GPUs
- 4K HDR and RAW output for post-production
- On-screen text rendering still unreliable
- Setup assumes solid GPU know-how
- Upscalers from version 2.3 must be chained separately
Multi-shot scenes, synchronized sound and complexity-aware rendering
LTX 2.5 generates video and its soundtrack in one go, with no separate audio stage. Native multi-shot generation defines this release, a single prompt describes several shots and the model keeps the character, set, lighting and voice steady from one cut to the next. You ask for a wide shot then a close-up, and the same actress returns with the same voice after the edit.
Under the hood, a diffusion video decoder replaces the old VAE reconstruction and drops visible glitches from 0.74 to 0.28 per clip on the internal 98-prompt test. A rendering approach the company calls Diffusion Fidelity Rendering spends compute where a scene gets hard. Native 4K HDR and RAW output feeds straight into video editing tools and professional color pipelines without conversion.
LTX 2.5 on your own hardware, from desktop GPU to server
A GPU with 16 GB of VRAM is enough to run LTX 2.5 at home. The model packs 22 billion parameters, backed by a Gemma text encoder, and the LTX family has passed 33 million downloads. Speed remains the headline act, 10 seconds of video come out in 6.8 seconds of compute (measured by LTX itself on two GB200 chips at 720p, worth saying upfront).
Against closed rivals on that same vendor-run benchmark, the gap speaks plainly.
| Model | Measured time for 10 s of video |
|---|---|
| LTX 2.5 self-hosted (2x GB200) | 6.8 s |
| LTX 2.5 via the API | 23.7 s |
| Omni Flash, Grok 1.5, Veo 3.1 | 52 to 70 s |
| Seedance 2.5 | 317 s |
| Kling 3.0 Pro | 398 s |
Getting started with LTX 2.5, from free weights to the per-second API
Three routes lead to the model, none of them gated by a waitlist.
On pricing, the Fast variant starts at $0.09 per second in 720p and climbs to $0.30 in 4K, while Pro sits at $0.12 in 720p and $0.17 in 1080p. The Community License makes commercial use free for organizations under $10 million in annual revenue. These rates shift between releases, so treat the ltx.io pricing page as the only figure worth budgeting against.
- Hugging Face and GitHub, with free weights, inference code and a fine-tuning checkpoint
- ComfyUI, with official workflow templates available from day one
- The LTX API, plus Runway and fal for a hosted route with zero installs
Frequently asked questions
Is LTX 2.5 free?
Yes, the weights, inference code and training code download free from Hugging Face and GitHub. The Community License covers commercial use as long as your company stays under $10 million in annual revenue. Only the API and third-party hosts charge, billed per second of generated video.
What GPU does LTX 2.5 require to run locally?
A 16 GB VRAM card is the floor, and the model is tuned for NVIDIA RTX GPUs and DGX Spark. For comfortable production work, even with the distilled checkpoint, something closer to 48 GB makes more sense. Without an NVIDIA GPU, the API or a hosted platform fills the gap.
LTX 2.5 or Veo 3.1, what is the difference?
Veo 3.1 is a closed model reachable only through Google's cloud, while LTX 2.5 downloads, fine-tunes and runs on your own machines. On LTX's published measurements, generation runs several times faster on LTX 2.5, with Veo keeping a strong reputation for image fidelity. Local control usually decides the matter.
What is the physical AI checkpoint in LTX 2.5 for?
Robot training is the target. A pretrained checkpoint tuned for physical AI ships alongside the main model, and robotics teams fine-tune it on their own footage so machines can predict how an environment reacts to their actions. Markov Robotics is among the first companies putting it to work.
Verdict: Ten seconds of footage in under seven seconds of compute turns video generation into a near-instant trial loop. GPU-equipped studios, ad teams cycling through product-shot variations and developers building on open weights will feel right at home.
