arXiv:2503.10342cs.CV2025-03被引 8

仅用一张图片实现物体零样本插入视频,无需训练或额外数据。

DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image

  • 通过分析物体运动轨迹预测未知动作,实现自然融合。
  • 无需训练或微调,直接在真实视频中插入新物体。
  • 适合影视制作、创意设计等内容生成场景。

生成式扩散模型让许多梦想成为现实。然而,现有视频物体插入方法通常需要参考视频或物体的3D资产以生成合成运动。从单张参考图像向目标背景视频中插入物体,因缺乏未见运动信息而仍属空白领域。本文提出DreamInsert,首次实现无需训练的图像到视频物体插入。通过考虑物体运动轨迹,DreamInsert能预测未知运动,与背景视频自然融合,生成无缝结果。更重要的是,该方法简单高效,无需端到端训练或在精心设计的图像-视频配对数据上微调,即可实现零样本插入。我们通过多种实验验证其有效性,并首次展示无训练条件下图像到视频物体插入的结果,为未来内容创作与合成开辟新方向。代码即将发布。

原文摘要 · Abstract (English)

Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, such as a reference video or a 3D asset of the object, to generate the synthetic motion. However, inserting an object from a single reference photo into a target background video remains an uncharted area due to the lack of unseen motion information. We propose DreamInsert, which achieves Image-to-Video Object Insertion in a training-free manner for the first time. By incorporating the trajectory of the object into consideration, DreamInsert can predict the unseen object movement, fuse it harmoniously with the background video, and generate the desired video seamlessly. More significantly, DreamInsert is both simple and effective, achieving zero-shot insertion without end-to-end training or additional fine-tuning on well-designed image-video data pairs. We demonstrated the effectiveness of DreamInsert through a variety of experiments. Leveraging this capability, we present the first results for Image-to-Video object insertion in a training-free manner, paving exciting new directions for future content creation and synthesis. The code will be released soon.

视频生成零样本扩散模型物体插入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。