LTX-2.5: The Open-Source Breakthrough for Multishot AI Videos
With LTX-2.5, Lightricks launches a groundbreaking open-weights model for AI videos. Learn all about native multishots, synchronized sound, and extreme rendering speeds.
Introduction
The world of generative video AI is evolving at a breakneck pace, but a new release is currently turning the entire creator scene upside down. With the release of LTX-2.5 (as of August 2026), the development team at LTX (spun off from the well-known software house Lightricks) has made an extremely powerful open-weights model available. The 22-billion-parameter model breaks with old limitations and proves that cinematic video creation no longer has to take place solely behind the paywalls of tech giants.
Until now, creators often had to struggle with extremely long rendering times or painstakingly piece together individual scenes in the hope that character and environmental consistency would be maintained. LTX-2.5 solves these problems in one fell swoop. It combines groundbreaking speed with an architecture that can compute native film cuts and lip-sync audio directly in a single generation run.
For anyone looking to get started directly on the web and create professional content without running their own high-end graphics cards, AIQORA offers direct access to state-of-the-art tools: simply use our AI Video Generator or enhance your projects with the matching Music Generator to create cinematic masterpieces in no time.
Key Takeaways
* Native Multishot Scenes: For the first time, LTX-2.5 generates coherent scenes with multiple cuts in a single run, keeping characters, lighting, and audio perfectly consistent across cuts.
* Near-Real-Time Rendering Speed: On high-performance hardware (such as 2× Nvidia GB200 GPUs), the model generates a 10-second video in just 6.8 seconds at 720p resolution.
* Integrated Audio Generator: Visuals and sound are calculated simultaneously—noises, speech, and atmosphere match the image perfectly from second one.
* Open-Weights License: The model is available for free on Hugging Face and is completely royalty-free for commercial use by organizations with annual revenues under $10 million.
* Gemma 4 12B Integration: Thanks to a specially modified text encoder, the model understands complex, multi-part camera and directing instructions flawlessly.
The Video AI Revolution: What Makes LTX-2.5 So Special?
While many conventional models are designed to generate a single, continuous video clip without sound, LTX-2.5 takes a holistic approach to what is known as "world modelling." It is optimized to deeply understand real physical processes and cinematic structures.
Native Multishot Generation
Anyone who has ever tried to create a short film using video AI knows the problem: you generate a wide shot of a person, then cut to a close-up, and suddenly the character has a different face or is wearing a different jacket. LTX-2.5 solves this problem through natively generated cuts. The model calculates a sequence with different camera angles in a single run. The consistency of character identity, clothing, scenery, and lighting mood remains absolutely stable throughout.
Synchronized Audio Directly from Diffusion
Instead of painstakingly creating and adjusting audio tracks afterwards using third-party tools, LTX-2.5 generates synchronous audio in the same computing process. If a wave breaks in the video or a character speaks, the corresponding sound is synchronized exactly to the frame. This saves creators an enormous amount of time during the sound design workflow.
Detail Fidelity with a New Decoder
Conventional VAE decoders (Variational Autoencoders) tend to make fine details like faces, textures, or lettering look muddy in the video. LTX-2.5 introduces a completely novel Diffusion Video Decoder for this purpose. This performs a second, fine diffusion process during image reconstruction, ensuring that faces remain razor-sharp and text in the video is legible. This is supported by a dynamic computing power distribution system (Diffusion Fidelity Rendering), which allocates more computing power to complex scene areas.
Technical Excellence: Gemma 4 and ComfyUI Integration
The control of video models often fails because long, detailed prompts are "forgotten" by the system. LTX-2.5 relies on a tailored Gemma 4 12B text encoder for this. This ensures that even complex directing instructions—such as "camera pans from left to right while a car drives by in the background and the main character looks up in surprise"—are implemented in detail and without loss of information.
For professional creators setting up their workflows locally, LTX-2.5 is excellently integrated into ComfyUI. It supports three primary workflows out-of-the-box:
- Text-to-Video (T2V): Creation of scenes purely from text descriptions.
- Image-to-Video (I2V): A starting image (keyframe) serves as a visual template for the video motion.
- Frame Interpolation (FLF2V): Generating motion precisely between a predefined start and end image.
LTX-2.5 vs. MiniMax H3: The Clash of Titans
In the current market environment (as of August 2026), LTX-2.5 is primarily competing with the MiniMax H3 model. Both pursue exciting open-weights approaches but show different strengths:
| Feature / Capability | LTX-2.5 (Lightricks) | MiniMax H3 |
|---|---|---|
| Focus | Maximum speed, native multishots, 4K HDR pipelines | Complex physical interactions, 3D camera movements |
| Audio | Synchronous audio (mono/ambient) integrated | Native 32 kHz stereo audio & multilingual dialogues |
| Speed | Extremely fast (distilled version with only 8 steps) | Moderate, more complex calculations |
| Workflow | Perfect for fast prototyping and native post-production (EXR export) | Strong in scenic, physically demanding single clips |
Conclusion: The Democratization of Video Production
With LTX-2.5, the open-source community impressively demonstrates that world-class video AI does not necessarily have to be tied to expensive cloud subscriptions. The combination of rapid rendering speeds, native scene continuity across cuts, and lip-sync audio makes the model an indispensable tool for filmmakers, advertising agencies, and content creators.
If you want to experience the power of modern AI creation directly in your browser without dealing with complex installations like ComfyUI, the AIQORA platform offers you the perfect alternative. Explore our intuitive AI Video Generator, create fitting soundtracks with the Music Generator, or generate stunning graphics with our Image Generator to turn your creative vision into reality!
FAQ
Is LTX-2.5 completely free to use?
Yes, the weights of the model (open-weights) can be downloaded for free on Hugging Face. There are no licensing fees for commercial use, provided the generating company or organization has an annual revenue of under $10 million.
What hardware do I need to run LTX-2.5 locally?
To run the full 22-billion-parameter model locally at an acceptable speed, a modern Nvidia graphics card with at least 16 GB VRAM (preferably 24 GB VRAM like an RTX 3090/4090 or higher) is recommended. For less powerful systems, there is an optimized "distilled" version that requires significantly fewer resources.
Can LTX-2.5 output videos in 4K resolution?
Yes, LTX-2.5 supports native resolutions from 720p up to 4K HDR. However, at extremely high resolutions like 4K, the maximum standard duration of a single clip is usually limited to 10 seconds to avoid overloading the graphics memory.
What does "Native Multishot Generation" mean?
This means that the model generates a finished video in a single calculation run that already contains cuts (e.g., a medium-wide shot followed by a close-up). The visual identity of the characters, the scenery, and the voices remain absolutely consistent across the cuts.