arXiv:2602.09609cs.CV2026-02被引 8

统一视频生成与编辑框架,支持多模态指令控制

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

  • 用预训练多模态大模型解析文本、图像、参考视频等指令
  • 支持文本/图像到视频生成及上下文视频编辑等多种任务
  • 解耦指令解析与视频合成,提升可控性与一致性

基于扩散模型的视频生成技术显著提升了视觉质量与时间连贯性。然而,现有方法多为特定任务设计,主要依赖文本指令,难以处理多模态输入、上下文引用及多样化的生成与编辑场景。许多视频编辑方法依赖针对单一操作定制的复杂流水线,限制了可扩展性与组合性。本文提出Tele-Omni,一个统一的多模态视频生成与编辑框架,可在单一模型中响应包括文本、图像和参考视频在内的多模态指令。Tele-Omni利用预训练多模态大语言模型解析异构指令并推断结构化生成或编辑意图,扩散模型则根据这些结构化信号进行高质量视频合成。为实现跨异构任务的联合训练,引入任务感知数据处理流程,将多模态输入统一为结构化指令格式,同时保留任务特异性约束。Tele-Omni支持多种视频中心任务,包括文本到视频生成、图像到视频生成、首尾帧视频生成、上下文视频生成及上下文视频编辑。通过解耦指令解析与视频合成,并结合任务感知数据设计,实现了灵活的多模态控制,同时保持强时间连贯性与视觉一致性。实验表明,Tele-Omni在多个任务上达到具有竞争力的性能。

原文摘要 · Abstract (English)

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their ability to handle multimodal inputs, contextual references, and diverse video generation and editing scenarios within a unified framework. Moreover, many video editing methods depend on carefully engineered pipelines tailored to individual operations, which hinders scalability and composability. In this paper, we propose Tele-Omni, a unified multimodal framework for video generation and editing that follows multimodal instructions, including text, images, and reference videos, within a single model. Tele-Omni leverages pretrained multimodal large language models to parse heterogeneous instructions and infer structured generation or editing intents, while diffusion-based generators perform high-quality video synthesis conditioned on these structured signals. To enable joint training across heterogeneous video tasks, we introduce a task-aware data processing pipeline that unifies multimodal inputs into a structured instruction format while preserving task-specific constraints. Tele-Omni supports a wide range of video-centric tasks, including text-to-video generation, image-to-video generation, first-last-frame video generation, in-context video generation, and in-context video editing. By decoupling instruction parsing from video synthesis and combining it with task-aware data design, Tele-Omni achieves flexible multimodal control while maintaining strong temporal coherence and visual consistency. Experimental results demonstrate that Tele-Omni achieves competitive performance across multiple tasks.

视频生成多模态扩散模型编辑框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。