arXiv:2605.24652cs.AIcs.CV2026-05被引 3

为音视频生成模型设计了一个全自动、以人为中心的评估基准。

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

论文配图:AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
图 1 · 摘自论文原文
  • 构建十维细粒度指标,覆盖视觉、音频与跨模态一致性。
  • 通过偏好学习训练专业评估器,可识别细微跨模态不一致。
  • 输出连续评分,更贴近人类判断,适用于强化学习优化。

音视频(AV)生成技术快速发展,实现了高保真且声画同步的人类相关场景合成,但评估仍处于初级阶段,现有基准多为粗粒度,依赖通用多模态大模型进行有限评估,难以准确衡量模型能力。为此,我们提出AVBench,一个专为人类相关音视频生成设计的全自动评估基准。其核心包含两项设计:(i) 以人为中心的细粒度指标。AVBench整合了十项针对真实人类场景的评估维度,涵盖视觉质量、音频质量及多层级跨模态一致性,捕捉现有基准常忽略的人类相关细节。(ii) 基于偏好学习的专用评估器。通过将真实视频转化为带受控扰动的多样化训练对,构建大规模监督数据;在高质量数据上微调后,评估器能可靠检测细微跨模态不一致。关键在于,评估不输出离散文本判断,而是从模型对二元决策的预测置信度中提取连续评分,该概率化评分机制比传统VQA式评估更可靠,且更接近人类判断。综合来看,AVBench提供自动化评估能力,具备强数据筛选潜力,并可作为强化学习中人类反馈(RLHF)的可微奖励信号。

原文摘要 · Abstract (English)

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).

音视频生成自动评估跨模态一致性偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。