arXiv:2602.00607cs.MMcs.SD2026-02ACL被引 8

针对多人对话生成视频的缺陷,构建了可精准诊断的评测基准。

MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation

  • 通过半自动流程生成1800段视频,设计2400个问答对进行细粒度故障分析。
  • 在四个层面评估生成效果:音画保真度、时间一致性、社交互动与电影表现力。
  • 适合想提升多说话人音视频生成质量的研究者和开发者使用。

近年来,文本到音视频(T2AV)生成技术已能合成多参与者对话的音视频内容。然而,现有评测基准大多针对真人录制视频或单说话人场景,难以有效诊断生成视频中的人设漂移、对话切换不自然、音画不同步等结构性问题。为此,我们提出MTAVG-Bench,一个面向多说话人对话中心音视频生成的故障驱动型诊断基准。该基准通过半自动流程生成1800段视频,采用精心设计的提示词,产出2400个人工标注的问答对,实现细粒度故障诊断。评估涵盖音画信号保真度、时间属性一致性、社交互动与电影表达四个层次。基于分层故障分类体系与定向问答协议,主要评估专有与开源全能模型识别多说话人T2AV输出故障模式的能力。我们在MTAVG-Bench上评测了12款专有与开源全能模型,其中Gemini 3 Pro整体表现最强,而主流开源模型在信号保真度与一致性方面仍具竞争力。总体而言,MTAVG-Bench支持精细化故障分析,推动模型对比与生成质量优化。

原文摘要 · Abstract (English)

Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchmarks remain largely designed for human-recorded videos or single-speaker settings. As a result, structural failures in generated multi-talker dialogue videos, such as identity drift, unnatural turn transitions, and audio-visual misalignment, cannot be effectively diagnosed. To address this issue, we introduce MTAVG-Bench, a failure-driven diagnostic benchmark for multi-talker dialogue-centric audio-video generation. MTAVG-Bench is built via a semi-automatic pipeline, where 1.8k videos are generated using mainstream T2AV models with carefully designed prompts, yielding 2.4k manually annotated QA pairs for fine-grained failure diagnosis. The benchmark evaluates multi-speaker dialogue generation at four levels: audio-visual signal fidelity, temporal attribute consistency, social interaction, and cinematic expression. Built on a hierarchical failure taxonomy and a targeted QA protocol, MTAVG-Bench is primarily designed to evaluate whether proprietary and open-source omni-models can reliably identify failure modes in multi-speaker T2AV outputs. We benchmark 12 proprietary and open-source omni-models on MTAVG-Bench, with Gemini 3 Pro achieving the strongest overall performance, while leading open-source models remain competitive in signal fidelity and consistency. Overall, MTAVG-Bench enables fine-grained failure analysis for rigorous model comparison and targeted video generation refinement.

音视频生成多说话人评测基准故障诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。