一句话生成高质量视频,支持图文指令与智能编辑。
Kling-Omni Technical Report
- 端到端框架融合视频生成、编辑与推理任务。
- 支持文本、图片、视频输入,生成电影级质量内容。
- 适合需要多模态智能创作的开发者与创作者。
我们提出Kling-Omni,一种通用生成框架,可直接从多模态视觉语言输入合成高保真视频。采用端到端设计,该框架弥合了视频生成、编辑与智能推理任务之间的功能分离,将其整合为一个整体系统。不同于分散的流水线方法,Kling-Omni支持多种用户输入,包括文本指令、参考图像和视频上下文,将其统一编码为多模态表示,以生成具有电影级质量与高度智能性的视频内容。为支撑这些能力,我们构建了全面的数据系统,作为多模态视频创作的基础。框架还通过高效的超大规模预训练策略与推理基础设施优化得到增强。综合评估表明,Kling-Omni在上下文生成、基于推理的编辑和多模态指令遵循方面表现出色。超越内容创作工具,我们认为Kling-Omni是迈向多模态世界模拟器的关键进展,具备感知、推理、生成与交互动态复杂世界的能力。
原文摘要 · Abstract (English)
We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional separation among diverse video generation, editing, and intelligent reasoning tasks, integrating them into a holistic system. Unlike disjointed pipeline approaches, Kling-Omni supports a diverse range of user inputs, including text instructions, reference images, and video contexts, processing them into a unified multimodal representation to deliver cinematic-quality and highly-intelligent video content creation. To support these capabilities, we constructed a comprehensive data system that serves as the foundation for multimodal video creation. The framework is further empowered by efficient large-scale pre-training strategies and infrastructure optimizations for inference. Comprehensive evaluations reveal that Kling-Omni demonstrates exceptional capabilities in in-context generation, reasoning-based editing, and multimodal instruction following. Moving beyond a content creation tool, we believe Kling-Omni is a pivotal advancement toward multimodal world simulators capable of perceiving, reasoning, generating and interacting with the dynamic and complex worlds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。