arXiv:2501.18314cs.MMcs.CV2025-01ICML被引 15

用大模型评估AI生成音视频质量,提升配音效果与用户体验。

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

  • 基于大视觉语言模型构建多维评分模型,支持音视频质量评估。
  • 在3382个AI生成音视频上达到领先性能,显著优于现有方法。
  • 适合音视频生成、质量评估与人机交互研究者使用。

许多视频到音频(VTA)方法被提出用于为无声AI生成视频配音。高效评估AI生成音视频(AGAV)质量的方法对保障音视频质量至关重要。现有音视频质量评估方法难以应对AGAV中独特的失真问题,如不真实和不一致的内容。为此,我们提出了AGAVQA-3k,首个大规模AGAV质量评估数据集,包含来自16种VTA方法的3,382个AGAV。该数据集包含两个子集:AGAVQA-MOS提供音频质量、内容一致性及整体质量的多维度评分;AGAVQA-Pair用于最优AGAV对选择。我们进一步提出AGAV-Rater,一种基于大模型的评分系统,可对AGAV、文本生成音频及音乐进行多维度评分,并从多个VTA生成结果中选出最佳方案呈现给用户。AGAV-Rater在AGAVQA-3k、Text-to-Audio和Text-to-Music数据集上均达当前最优表现。主观测试也证实其能有效提升VTA性能与用户体验。数据集与代码已开源:https://github.com/charlotte9524/AGAV-Rater。

原文摘要 · Abstract (English)

Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is crucial for ensuring audio-visual quality. Existing audio-visual quality assessment methods struggle with unique distortions in AGAVs, such as unrealistic and inconsistent elements. To address this, we introduce AGAVQA-3k, the first large-scale AGAV quality assessment dataset, comprising $3,382$ AGAVs from $16$ VTA methods. AGAVQA-3k includes two subsets: AGAVQA-MOS, which provides multi-dimensional scores for audio quality, content consistency, and overall quality, and AGAVQA-Pair, designed for optimal AGAV pair selection. We further propose AGAV-Rater, a LMM-based model that can score AGAVs, as well as audio and music generated from text, across multiple dimensions, and selects the best AGAV generated by VTA methods to present to the user. AGAV-Rater achieves state-of-the-art performance on AGAVQA-3k, Text-to-Audio, and Text-to-Music datasets. Subjective tests also confirm that AGAV-Rater enhances VTA performance and user experience. The dataset and code is available at https://github.com/charlotte9524/AGAV-Rater.

音视频生成质量评估大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。