Kuaishou's Kling AI has unveiled Kling 4.0, its latest flagship video generation model, now available in preview. The update represents one of the largest capability jumps in the model's history, addressing key limitations in clip length, reference management, resolution, and audio integration. For businesses and creators relying on AI-generated video, these improvements could significantly reduce post-production work and expand creative possibilities.
At the core of the update is extended clip generation. Kling 4.0 Preview supports native generation of up to 30 seconds per clip, while an experimental Long Video mode can produce continuous 120-second sequences at 1080p. Clips can also be extended to 60 seconds after generation. This directly responds to a frequent complaint about earlier Kling versions, where shorter native runtimes made narrative and advertising projects harder to plan. Longer native clips mean fewer edits and more coherent storytelling, which is crucial for marketing campaigns and short films.
The centerpiece is Omni Reference, a system that allows a single generation to draw on up to 15 reference elements from a pool of as many as 50 uploaded files—30 images, 10 video clips, and 10 audio clips. Creators can lock a character's face across multiple shots, match a product's exact color from a reference photo, or carry a specific lighting setup from one scene to the next without re-describing those details in every prompt. This level of consistency is a major step forward for professional workflows, where maintaining visual continuity is essential.
Multi-shot generation builds on the same reference system. A single prompt can now produce a sequence of shots that hold spatial continuity—a room stays the same room as the camera moves, and characters keep the same clothing and face across cuts. Historically, AI video tools often visibly 'reset' a scene when the camera angle changed, breaking immersion. By solving this, Kling 4.0 enables more complex narratives and reduces the need for manual fixes.
On the audio front, dialogue, ambient sound, and music are generated in the same pass as the picture, rather than layered on afterward. Lip movement is synced to spoken lines across multiple languages and accents, and different lines can be assigned to different characters within the same scene. This removes a manual dubbing step that previously required separate audio software, streamlining production and opening doors for multilingual content creation.
Native 4K output (3840×2160) rounds out the release, aimed at preserving fine texture, sharp edges, and depth of field through fast-motion shots rather than relying on post-generation upscaling. Creators evaluating AI video tools have increasingly flagged native 4K as the real test of whether a clip survives the move from preview window to actual publish. For industries like advertising, film, and e-learning, this means higher-quality output straight from the generator.
Kling4.org provides direct access to Kling 4.0's generation modes as they roll out—including standard, Pro, and native 4K tiers—letting creators, marketers, and small teams begin testing prompt-to-clip workflows, persistent-character generation, and multi-shot sequencing. The site also offers a Kling 4.0 video prompt library for creators who want practical starting points for their own scenes. As AI video continues to evolve, Kling 4.0's advancements could lower barriers to entry, enabling smaller teams to produce professional-grade content without extensive resources.

