arXiv:2501.01427cs.CV2025-01International Conf…被引 47

零样本视频物体插入,细节清晰且运动精准可控。

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

  • 用身份提取器注入全局身份,用框序列控制整体运动
  • 设计像素扭曲模块,实现细节保真与精细轨迹操控
  • 支持换脸、虚拟试穿等应用,无需额外微调

尽管视频生成技术取得显著进展,将指定物体插入视频仍具挑战性,核心难点在于同时保持参考物体的外观细节并准确建模连贯运动。本文提出 VideoAnydoor,一种零样本视频物体插入框架,兼具高保真细节保留与精确运动控制能力。基于文本到视频模型,通过身份提取器注入全局身份,并利用框序列控制整体运动。为在保持细节的同时支持细粒度运动控制,设计像素扭曲模块:输入参考图像及任意关键点与对应轨迹,根据轨迹扭曲像素细节,并融合至扩散U-Net,提升细节保留并支持用户操控运动轨迹。此外,提出结合视频与静态图像的训练策略,采用加权损失增强插入质量。VideoAnydoor 在多项指标上显著优于现有方法,自然支持多种下游应用(如说话头生成、视频虚拟试穿、多区域编辑),无需任务特定微调。

原文摘要 · Abstract (English)

Despite significant advancements in video generation, inserting a given object into videos remains a challenging task. The difficulty lies in preserving the appearance details of the reference object and accurately modeling coherent motions at the same time. In this paper, we propose VideoAnydoor, a zero-shot video object insertion framework with high-fidelity detail preservation and precise motion control. Starting from a text-to-video model, we utilize an ID extractor to inject the global identity and leverage a box sequence to control the overall motion. To preserve the detailed appearance and meanwhile support fine-grained motion control, we design a pixel warper. It takes the reference image with arbitrary key-points and the corresponding key-point trajectories as inputs. It warps the pixel details according to the trajectories and fuses the warped features with the diffusion U-Net, thus improving detail preservation and supporting users in manipulating the motion trajectories. In addition, we propose a training strategy involving both videos and static images with a weighted loss to enhance insertion quality. VideoAnydoor demonstrates significant superiority over existing methods and naturally supports various downstream applications (e.g., talking head generation, video virtual try-on, multi-region editing) without task-specific fine-tuning.

视频生成物体插入运动控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。