arXiv:2605.28203cs.LG2026-05

针对视频生成评估中多维度监督风险不均的问题,提出解耦影响函数方法提升奖励模型精度。

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

论文配图:Refining Multidimensional Video Reward Models via Disentangled Influence Functions
图 1 · 摘自论文原文
  • 基于解耦影响函数,分别评估每维监督风险,避免全局筛选的误判。
  • 在多个数据集上,新方法使奖励模型与真实评分相关性提升12.3%以上。
  • 适合需要精细视频质量评估的生成模型训练者使用。

随着文本到视频生成模型的发展,视频评估需在多个维度上进行细粒度判断。现有多维视频奖励模型(MVRMs)虽能分解评估任务以匹配人类视觉感知的多面性,但其训练受视频数据复杂性制约。本文发现关键现象——维度异质性:同一样本在不同评价维度上的可靠性差异显著,可能对某一目标提供可靠监督,却对另一目标带来高风险。因此,依赖全局标量指标的数据筛选方法不适用于文本到视频任务。为此,我们提出解耦影响函数框架,高效估计各维度的监督风险。基于此,设计两种解耦优化策略:维度解耦剪枝,移除极端高风险样本;维度解耦重加权,软性降低高风险监督权重。大量实验表明,该方法显著优于全局过滤基线,在多个数据集上实现更优的真实评分对齐效果。

原文摘要 · Abstract (English)

As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various axes. To address this, recent works have focused on developing Multidimensional Video Reward Models (MVRMs), which decompose the evaluation process to better align with the multifaceted nature of human visual perception. However, training effective MVRMs is fundamentally challenged by the complex nature of video data. In this work, we identify a critical phenomenon termed Dimensional Heterogeneity: the reliability of a training sample can vary substantially across evaluation dimensions, meaning that a sample may provide reliable supervision for one objective while inducing high supervision risk for another. Consequently, prevailing data-centric methods that filter based on global scalar metrics are ill-posed for T2V tasks. To address this, we propose a disentangled influence framework that that efficiently estimates dimension-specific supervision risk. Leveraging this framework, we introduce two dimension-disentangled refinement strategies: Dimension-Disentangled Pruning, which removes extreme high-risk samples, and Dimension-Disentangled Reweighting, which softly down-weights high-risk supervision. Extensive experiments demonstrate that our disentangled strategies significantly outperform global filtering baselines, yielding reward models with superior alignment to ground truth.

视频生成奖励模型多维评估解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。