LTX, the open world model company spun out of Lightricks, has released LTX-2.5 — a rebuilt video and world model that generates a 10-second, 720p clip from an image in 6.8 seconds when self-hosted on two NVIDIA GB200 superchips. The model is available now as open weights on Hugging Face, natively integrated into ComfyUI through a day-one partnership, and via a managed API. Organizations under $10 million in annual recurring revenue can use it free; larger companies negotiate a license.
The release marks a deep pipeline overhaul. LTX-2.5 introduces a new diffusion video decoder that reduces artifacts in high-motion footage while preserving fine detail like text and faces. Native multishot generation renders a full sequence in a single pass, holding character, scene, and voice consistent across cuts. A custom Gemma 4 language backbone and dedicated prompt enhancer improve handling of complex, multi-subject prompts. A pretrained checkpoint tuned for physical AI and robotics gives teams a base to fine-tune on domain data that looks nothing like cinematic video. An optimized distilled model runs locally on NVIDIA RTX GPUs with a 16 GB VRAM minimum.
What's new
- Diffusion video decoder: Reconstructs fine detail and reduces visual artifacts in high-motion footage while keeping LTX's high compression ratio.
- Native multishot generation: Renders a full sequence as a single output, maintaining consistency across cuts instead of stitching individually generated shots.
- Custom Gemma 4 backbone: Language model paired with a prompt enhancer for more accurate multi-subject prompt adherence.
- Robotics checkpoint: Pretrained weights tuned for physical AI and robotics, enabling fine-tuning on non-cinematic domain data.
- Distilled model: Near-full-model quality at lower cost and faster inference, optimized with NVIDIA to run on RTX GPUs and Macs.
- ComfyUI native integration: Day-one strategic partnership makes LTX-2.5 a first-class node in the de facto prototyping environment for open generative media.
On pricing, LTX-2.5 Fast tier generates 720p video with audio at $0.09 per second — $0.90 for a 10-second clip. That positions it between Google's Veo 3.1 Lite ($0.50 per 10 seconds) and Veo 3.1 Fast ($1.00), while undercutting FLUX 3 Video ($1.70) and HappyHorse 1.0 (~$1.82). The Pro tier runs at $0.12 per second ($1.20 per clip) with quality tuning for prompt adherence, faces, and typography, topping out at 1080p and 10 seconds. Self-hosting the open weights eliminates per-generation costs entirely for qualifying organizations.
Speed claims come with hardware caveats. The headline 6.8-second figure was measured self-hosted on two GB200 chips at steady state — a configuration far beyond most teams. Through LTX's managed API, the same job took 23.7 seconds at 1080p (the API has no 720p tier). By LTX's own end-to-end measurements of competing APIs, Gemini Omni Flash came in at 52 seconds, Grok 1.5 at 63 seconds, Veo 3.1 at 70 seconds for an 8-second clip, MiniMax H3 at 180 seconds, Seedance 2.5 at 317 seconds, and Kling 3.0 Pro at 398 seconds.
Quality claims are similarly vendor-reported. In blind, side-by-side human preference tests commissioned by LTX, LTX-2.5 recorded a 67% win rate, narrowly ahead of Seedance 2.5 at 65%, with Gemini Omni Flash at 55%, MiniMax H3 at 50%, and FLUX 3 at 28%. The company labels these results preliminary and expects them to evolve. The independent Artificial Analysis text-to-video arena leaderboards currently place Gemini Omni Flash at the top and do not yet score LTX-2.5.
Why it matters
The release reinforces a business model bet that open weights — not closed APIs — will win the video and world model market. CEO Zeev Farbman argues that video and world models have a fundamentally wider "surface area" of use cases than language models, requiring builders to access weights and create custom flows. The LTX-2.x Community License permits free use, modification, self-hosting, and sublicensing for organizations under the $10 million ARR threshold (measured across affiliates and subsidiaries). Even companies above the line can download and evaluate free in non-production environments.
The license's derivative definition is expansive: it covers fine-tuned checkpoints, LoRA adapters, distillations, and any model trained on LTX-2.5's outputs or synthetic data. All derivatives must be redistributed under the same license, and a fine-tune transferred to a company above the revenue threshold triggers that company's paid-license obligation. The license would not qualify as open source under the Open Source Initiative's definition due to revenue and field-of-use restrictions.
For robotics and physical AI teams, the dedicated checkpoint and local deployment capability matter. The model runs on-premises, at the edge, or via API with no visible watermark — though the license requires disclosure that content is machine-generated and forbids removing embedded provenance features. LTX says its model family has passed 33 million downloads, making it the most-used open world model line on the market.
Sources
- LTX-2.5 release announcement: 10-second AI video in 6.8 seconds on Nvidia superchips, open weights
- CNET: This New Open-Weight AI Model Is Built for Video and Robots
- Hugging Face: Lightricks/LTX-2.5 model repository
- AIbase: LTX Officially Launches Open World Model LTX-2.5
- LTX Blog: How To Generate 20 Second AI Videos With LTX-2.3