arXiv:2511.22974cs.CV2025-11被引 1

提出新方法让视频生成更符合人类偏好,尤其改善高动态画面质量。

McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning

  • 分维度解析人类偏好,用自省式推理训练奖励模型。
  • 通过分层对比推理,实现对视频整体质量的多维评估。
  • 动态调整优化目标,避免模型偏向低运动内容,适合高质量视频生成研究者。

文本到视频(T2V)生成在对齐文本提示方面取得了显著进展,但如何精准匹配人类细微偏好仍具挑战,因人类判断具有主观性和多维度特征。现有方法依赖昂贵的人工标注或使用代理指标预测偏好,缺乏对偏好逻辑的理解。同时,它们通常直接对齐模型与整体偏好分布,忽略了运动动态与视觉质量等潜在冲突维度,导致模型偏向低运动内容。为此,我们提出运动修正型自省分层推理偏好对齐框架(McSc),一个三阶段强化学习框架,实现稳健的偏好建模与对齐。首先,自省分维度推理(ScDR)训练生成式奖励模型(RM),通过自省推理链实现可靠的部分维度评估;其次,引入分层对比推理(HCR),结合分层奖励监督,进行结构化多维推理;最后,利用RM优选视频,提出运动修正直接偏好优化(McDPO),动态重加权对齐目标,缓解对低运动内容的偏差。实验表明,McSc在人类偏好对齐上表现更优,并生成更具高运动动态的视频。

原文摘要 · Abstract (English)

Text-to-video (T2V) generation has achieved remarkable progress in producing high-quality videos aligned with textual prompts. However, aligning synthesized videos with nuanced human preference remains challenging due to the subjective and multifaceted nature of human judgment. Existing video preference alignment methods rely on costly human annotations or utilize proxy metrics to predict preference, which lacks the understanding of human preference logic. Moreover, they usually directly align T2V models with the overall preference distribution, ignoring potential conflict dimensions like motion dynamics and visual quality, which may bias models towards low-motion content. To address these issues, we present Motion-corrective alignment with Self-critic hierarchical Reasoning (McSc), a three-stage reinforcement learning framework for robust preference modeling and alignment. Firstly, Self-critic Dimensional Reasoning (ScDR) trains a generative reward model (RM) to decompose preferences into per-dimension assessments, using self-critic reasoning chains for reliable learning. Secondly, to achieve holistic video comparison, we introduce Hierarchical Comparative Reasoning (HCR) for structural multi-dimensional reasoning with hierarchical reward supervision. Finally, using RM-preferred videos, we propose Motion-corrective Direct Preference Optimization (McDPO) to optimize T2V models, while dynamically re-weighting alignment objective to mitigate bias towards low-motion content. Experiments show that McSc achieves superior performance in human preference alignment and generates videos with high-motion dynamic.

视频生成偏好对齐强化学习运动动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。