arXiv:2510.14955cs.CVcs.AI2025-10被引 6

用真实视频提升生成动作的自然度,让虚拟动作更像真人。

RealDPO: Real or Not Real, that is the Preference

  • 用真实视频做正样本,通过偏好优化改进动作生成
  • 在多个指标上优于当前最强模型,动作更流畅真实
  • 适合需要高质量动作生成的研究与应用

视频生成模型在合成质量上取得显著进展,但复杂动作生成仍是关键挑战,现有模型常难以生成自然、平滑且上下文一致的动作。这一与真实动作的差距限制了其实际应用。为此,我们提出RealDPO,一种利用真实世界数据作为正样本进行偏好学习的新对齐范式,以提升动作真实性。不同于传统监督微调(SFT)提供的有限修正反馈,RealDPO采用定制损失函数的直接偏好优化(DPO),通过对比真实视频与错误模型输出,实现迭代自纠错,逐步提升动作质量。为支持复杂动作生成的后训练,我们构建了RealAction-5K,一个精选的高质量视频数据集,涵盖丰富精确的人类日常活动动作细节。大量实验表明,相较于最先进模型和现有偏好优化方法,RealDPO显著提升了视频质量、文本对齐度及动作真实性。

原文摘要 · Abstract (English)

Video generative models have recently achieved notable advancements in synthesis quality. However, generating complex motions remains a critical challenge, as existing models often struggle to produce natural, smooth, and contextually consistent movements. This gap between generated and real-world motions limits their practical applicability. To address this issue, we introduce RealDPO, a novel alignment paradigm that leverages real-world data as positive samples for preference learning, enabling more accurate motion synthesis. Unlike traditional supervised fine-tuning (SFT), which offers limited corrective feedback, RealDPO employs Direct Preference Optimization (DPO) with a tailored loss function to enhance motion realism. By contrasting real-world videos with erroneous model outputs, RealDPO enables iterative self-correction, progressively refining motion quality. To support post-training in complex motion synthesis, we propose RealAction-5K, a curated dataset of high-quality videos capturing human daily activities with rich and precise motion details. Extensive experiments demonstrate that RealDPO significantly improves video quality, text alignment, and motion realism compared to state-of-the-art models and existing preference optimization techniques.

视频生成动作生成偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。