arXiv:2508.10858cs.CV2025-08NeurIPS被引 22

通过分层偏好优化提升视频物理合理性,让生成视频更真实可信。

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

  • 分四个层级(实例、状态、运动、语义)进行细粒度对齐,增强物理一致性。
  • 在多个基准测试中显著提升视频的物理合理性与整体质量。
  • 自动筛选高质量数据,无需人工构建数据集,适合追求真实感的视频生成研究者。

近期视频生成技术已能生成高质量、视觉吸引人的视频,但确保生成视频符合物理规律仍是关键挑战。本文提出 PhysHPO 框架,一种分层跨模态直接偏好优化方法,实现细粒度物理合理性对齐。该框架在四个层级上优化:(a) 实例级,使视频内容与输入提示一致;(b) 状态级,以边界帧为锚点保障时序一致性;(c) 运动级,建模真实运动轨迹;(d) 语义级,保持叙事与视觉逻辑一致。鉴于真实视频是物理现象的最佳反映,我们设计自动化数据筛选流程,高效从大规模文本-视频数据集中提取优质样本,避免耗时费力的数据构建。在聚焦物理和通用能力的多项基准测试中,PhysHPO 显著提升先进模型的物理合理性与整体生成质量。据我们所知,这是首个探索视频生成中细粒度偏好对齐与数据选择的工作,为更真实、更受人类偏好的视频生成范式铺平道路。

原文摘要 · Abstract (English)

Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose PhysHPO, a novel framework for Hierarchical Cross-Modal Direct Preference Optimization, to tackle this challenge by enabling fine-grained preference alignment for physically plausible video generation. PhysHPO optimizes video alignment across four hierarchical granularities: a) Instance Level, aligning the overall video content with the input prompt; b) State Level, ensuring temporal consistency using boundary frames as anchors; c) Motion Level, modeling motion trajectories for realistic dynamics; and d) Semantic Level, maintaining logical consistency between narrative and visuals. Recognizing that real-world videos are the best reflections of physical phenomena, we further introduce an automated data selection pipeline to efficiently identify and utilize "good data" from existing large-scale text-video datasets, thereby eliminating the need for costly and time-intensive dataset construction. Extensive experiments on both physics-focused and general capability benchmarks demonstrate that PhysHPO significantly improves physical plausibility and overall video generation quality of advanced models. To the best of our knowledge, this is the first work to explore fine-grained preference alignment and data selection for video generation, paving the way for more realistic and human-preferred video generation paradigms.

视频生成物理合理性偏好优化分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。