arXiv:2509.20360cs.CV2025-09被引 51

统一图像视频编辑生成,用上下文学习实现跨模态自由操作

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

  • 将文本/图像/视频统一为序列令牌,通过自注意力实现上下文学习
  • 构建23.2万条视频编辑数据,联合训练提升多模态生成能力
  • 首个指令式视频编辑基准,适合需要灵活创作的开发者与创作者

近期基础模型的发展表明,统一与扩展是显著趋势,展现出跨领域的涌现能力。尽管图像生成与编辑已从专用任务转向统一框架,视频生成与编辑仍因架构限制和数据稀缺而分散。本文提出EditVerse,一个在单个模型中统一处理图像与视频生成及编辑的框架。通过将文本、图像、视频统一表示为序列令牌,EditVerse利用自注意力机制实现强大的上下文学习、自然的跨模态知识迁移,并灵活处理任意分辨率与时长的输入输出。为解决视频编辑训练数据不足问题,设计可扩展的数据流水线,收集23.2万条视频编辑样本,并与大规模图像和视频数据集联合训练。此外,提出EditVerseBench,首个基于指令的视频编辑基准,涵盖多样化任务与分辨率。大量实验与用户研究显示,EditVerse达到领先性能,超越现有开源与商业模型,并展现出跨模态的涌现编辑与生成能力。

原文摘要 · Abstract (English)

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due to architectural limitations and data scarcity. In this work, we introduce EditVerse, a unified framework for image and video generation and editing within a single model. By representing all modalities, i.e., text, image, and video, as a unified token sequence, EditVerse leverages self-attention to achieve robust in-context learning, natural cross-modal knowledge transfer, and flexible handling of inputs and outputs with arbitrary resolutions and durations. To address the lack of video editing training data, we design a scalable data pipeline that curates 232K video editing samples and combines them with large-scale image and video datasets for joint training. Furthermore, we present EditVerseBench, the first benchmark for instruction-based video editing covering diverse tasks and resolutions. Extensive experiments and user studies demonstrate that EditVerse achieves state-of-the-art performance, surpassing existing open-source and commercial models, while exhibiting emergent editing and generation abilities across modalities.

视频生成统一框架上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。