arXiv:2603.11089cs.SDcs.MM2026-03中稿 · ICASSP2026

用偏好优化提升视频转音频模型的听感质量与时间对齐

V2A-DPO: Omni-Preference Optimization for Video-to-Audio Generation

  • 基于流模型设计偏好优化框架,适配视频转音频任务
  • 在VGGSound上生成音频比基线模型更符合人类偏好
  • 适合关注音频生成质量与真实感的研究者

本文提出V2A-DPO,一种专为基于流的视频转音频(V2A)模型设计的直接偏好优化(DPO)框架,通过三项核心改进实现与人类偏好的对齐。首先引入AudioScore——一个用于评估合成音频语义一致性、时间对齐性和感知质量的人类偏好对齐评分系统;其次构建基于AudioScore的自动化流水线,生成大规模偏好对齐数据用于DPO优化;最后采用课程学习增强的DPO策略,专门适配流式生成模型。在基准VGGSound数据集上的实验表明,经由V2A-DPO优化的Frieren和MMAudio模型,在人类偏好度量上优于使用去噪扩散策略优化(DDPO)及预训练基线的版本。此外,DPO优化后的MMAudio在多个指标上达到当前最优性能,超越已发表的V2A模型。

原文摘要 · Abstract (English)

This paper introduces V2A-DPO, a novel Direct Preference Optimization (DPO) framework tailored for flow-based video-to-audio generation (V2A) models, incorporating key adaptations to effectively align generated audio with human preferences. Our approach incorporates three core innovations: (1) AudioScore-a comprehensive human preference-aligned scoring system for assessing semantic consistency, temporal alignment, and perceptual quality of synthesized audio; (2) an automated AudioScore-driven pipeline for generating large-scale preference pair data for DPO optimization; (3) a curriculum learning-empowered DPO optimization strategy specifically tailored for flow-based generative models. Experiments on benchmark VGGSound dataset demonstrate that human-preference aligned Frieren and MMAudio using V2A-DPO outperform their counterparts optimized using Denoising Diffusion Policy Optimization (DDPO) as well as pre-trained baselines. Furthermore, our DPO-optimized MMAudio achieves state-of-the-art performance across multiple metrics, surpassing published V2A models.

视频转音频偏好优化流模型音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。