arXiv:2504.06672cs.CV2025-04被引 5

用检索视频增强生成动作,让动画更自然真实。

RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism

  • 生成时检索相似视频作为动作参考,引导模型学习真实运动
  • 在多个基准上提升动作真实度,减少僵硬与不自然
  • 适用于任意文生视频模型,无需大量微调

视频生成正快速发展,得益于扩散模型的进步和更大更优数据集的出现。然而,由于数据维度高、任务复杂,生成高质量视频仍具挑战。现有研究多关注视觉质量和时间一致性(如闪烁问题),但生成视频在动作复杂性和物理合理性方面仍不足,常呈现静态或不合理运动。本文提出一种新框架,通过在生成阶段引入检索机制来提升动作真实性。检索到的视频作为语义锚点,为模型提供物体运动的真实示范。该流程可适配任意文本到视频扩散模型,仅需少量微调即可实现。我们在多个基准上验证了方法优势,包括主流评估指标和新提出的评测标准,同时展示其在其他场景的应用潜力。

原文摘要 · Abstract (English)

Video generation is experiencing rapid growth, driven by advances in diffusion models and the development of better and larger datasets. However, producing high-quality videos remains challenging due to the high-dimensional data and the complexity of the task. Recent efforts have primarily focused on enhancing visual quality and addressing temporal inconsistencies, such as flickering. Despite progress in these areas, the generated videos often fall short in terms of motion complexity and physical plausibility, with many outputs either appearing static or exhibiting unrealistic motion. In this work, we propose a framework to improve the realism of motion in generated videos, exploring a complementary direction to much of the existing literature. Specifically, we advocate for the incorporation of a retrieval mechanism during the generation phase. The retrieved videos act as grounding signals, providing the model with demonstrations of how the objects move. Our pipeline is designed to apply to any text-to-video diffusion model, conditioning a pretrained model on the retrieved samples with minimal fine-tuning. We demonstrate the superiority of our approach through established metrics, recently proposed benchmarks, and qualitative results, and we highlight additional applications of the framework.

视频生成扩散模型动作真实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。