arXiv:2503.06268cs.CV2025-03被引 13

让用户把特定实物精准融入视频,实现自然互动。

Get In Video: Add Anything You Want to the Video

  • 通过参考图生成视频编辑对,解决实例融合难题。
  • 提出新框架,保持对象身份与时空一致性。
  • 构建首个专项评测基准,推动个性化视频编辑发展。

视频编辑日益需要将特定现实物体融入现有画面,但现有方法难以捕捉目标的独特视觉特征,也难以确保实例与场景的自然交互。本文提出“Get-In-Video 编辑”这一新范式,用户提供参考图像以精确指定希望插入视频的视觉元素。针对训练数据稀缺和时空一致性维护的技术挑战,我们做出三项贡献:首先,构建 GetIn-1M 数据集,基于自动化的识别-跟踪-擦除流程(包含视频描述、显著实例识别、目标检测、时序跟踪与实例移除),生成高质量带完整标注(参考图像、追踪掩码、实例提示)的视频编辑对;其次,提出 GetInVideo 框架,采用具备3D全注意力机制的扩散变换器架构,可同步处理参考图像、条件视频与掩码,有效维持时间连贯性、保留视觉身份并实现自然场景交互;最后,建立 GetInBench 基准,首次系统评估该场景下的编辑性能。实验表明本方法在多项指标上显著优于基线,极大提升特定实体融入视频的质量与可行性。

原文摘要 · Abstract (English)

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure natural instance/scene interactions. We formalize this overlooked yet critical editing paradigm as "Get-In-Video Editing", where users provide reference images to precisely specify visual elements they wish to incorporate into videos. Addressing this task's dual challenges, severe training data scarcity and technical challenges in maintaining spatiotemporal coherence, we introduce three key contributions. First, we develop GetIn-1M dataset created through our automated Recognize-Track-Erase pipeline, which sequentially performs video captioning, salient instance identification, object detection, temporal tracking, and instance removal to generate high-quality video editing pairs with comprehensive annotations (reference image, tracking mask, instance prompt). Second, we present GetInVideo, a novel end-to-end framework that leverages a diffusion transformer architecture with 3D full attention to process reference images, condition videos, and masks simultaneously, maintaining temporal coherence, preserving visual identity, and ensuring natural scene interactions when integrating reference objects into videos. Finally, we establish GetInBench, the first comprehensive benchmark for Get-In-Video Editing scenario, demonstrating our approach's superior performance through extensive evaluations. Our work enables accessible, high-quality incorporation of specific real-world subjects into videos, significantly advancing personalized video editing capabilities.

视频编辑实例融合扩散模型3D注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。