arXiv:2608.11913cs.CV2026-08IJCV

用偏好优化提升视频生成音频的同步与质量

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

论文配图:HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
图 1 · 摘自论文原文
  • 双路视频表征保留时序动态与细节
  • 在线偏好优化使音频更符合人感知
  • 推理时自适应优化提升音视频对齐

视频到音频(V2A)生成面临时间同步精度低和感知质量差的挑战,主要源于视觉与听觉线索间复杂模糊的关系。现有方法通常将视频压缩为单一特征表示,导致时序动态和细粒度视觉信息大量丢失;且依赖重建目标训练,与人类对音频质量与适配性的判断关联弱。本文提出HarmoniDPO,一种融合偏好优化的扩散模型框架:(1) 采用全局上下文与帧级特征结合的双路视频表示,保留时序动态与语义细节;(2) 受强化学习中人类反馈(RLHF)启发,引入在线直接偏好优化(online-DPO),基于偏好判断微调扩散模型,提升音频感知质量与对齐度;(3) 提出双尺度扩散搜索(DDS),一种测试时缩放算法,可在推理阶段自适应优化输出保真度。实验表明,HarmoniDPO在音视频同步性和主观音频质量上优于当前最优方法,为从视频生成逼真、受人类青睐的音频提供了稳健解决方案。

原文摘要 · Abstract (English)

Video-to-audio (V2A) generation faces significant challenges in achieving precise temporal synchronization and high perceptual quality due to the complex, ambiguous relationship between visual and auditory cues. Existing methods typically compress video inputs into single feature representations, leading to significant loss of temporal dynamics and fine-grained visual information. These approaches also rely on reconstruction-based training objectives that poorly correlate with human perceptual judgments of audio quality and appropriateness. We propose HarmoniDPO, a novel framework that integrates preference-based optimization into diffusion-based V2A generation to address these limitations. (1) Our approach leverages a dual video representation: combining global context with frame-wise features to preserve temporal dynamics and semantic detail. (2) Inspired by reinforcement learning from human feedback (RLHF), HarmoniDPO employs online Direct Preference Optimization (online-DPO) to fine-tune a diffusion-based V2A model from preference judgments, enhancing perceptual quality and alignment. (3) Additionally, we introduce Dual-scale Diffusion Search (DDS), a test time scaling algorithm that adaptively optimizes output fidelity during inference. Experiments demonstrate that HarmoniDPO outperforms state-of-the-art methods in audio-video synchronization and subjective audio quality, offering a robust solution for generating realistic, human-preferred audio from video.

视频生成扩散模型音频生成偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。