arXiv:2605.28035cs.AIcs.MM2026-05被引 2

构建新基准,诊断多角色影视生成中的表演失误。

MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation

论文配图:MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
图 1 · 摘自论文原文
  • 提出涵盖表演、叙事等四维度的失败分类体系。
  • 构建超1万条问答评估样本,支持场景级故障定位。
  • 揭示当前大模型在复杂影视生成缺陷诊断中仍存短板。

近年来,多说话人音视频生成(MTAVG)模型在唇同步和视听对齐等基础指标上表现良好,但这些指标难以评估场景级生成中的电影化表现力。在多角色场景中,生成模型需超越视听真实感,传达连贯的角色表演与更高层次的电影特质。为此,本文提出MTAVG-Bench 2.0,一个用于诊断多说话人音视频生成中电影化表现力失效模式的基准。该基准聚焦短剧和场景级生成,建立涵盖表演、叙事、氛围与视听语言的高层级失败分类体系。基于此,构建超过1万条问答评估实例,包含短剧级评估子集与故障时间定位子集,系统评估通用大语言模型诊断高级视听故障的能力。实验表明,商用通用模型如Gemini显著优于其他评估者,但即使最强模型在复杂故障诊断中仍表现不足。结果表明,MTAVG-Bench 2.0为电影级多说话人音视频生成的故障诊断提供了系统性评测框架。

原文摘要 · Abstract (English)

In recent years, Multi-Talker Audio-Video Generation (MTAVG) models have shown promising performance on fundamental metrics such as lip-sync and audio-visual alignment. However, these metrics remain insufficient for assessing cinematic expressiveness in scene-level generation. In multi-character scenes, generation models must go beyond audio-visual realism to convey coherent character performance and other higher-level cinematic qualities. To fill this gap, we introduce MTAVG-Bench 2.0, a benchmark for diagnosing failure modes of cinematic expressiveness in multi-talker audio-video generation. Unlike prior settings that mainly focus on the quality of basic multi-turn dialogue, MTAVG-Bench 2.0 targets short-drama and scene-level generation, and establishes a high-level failure taxonomy spanning acting, narrative, atmosphere, and audio-visual language. Based on this taxonomy, we construct more than 10,000 question-answering evaluation instances, together with subsets for short-drama-level assessment and temporal localization of failure modes, to systematically evaluate the ability of omni large language models to diagnose high-level audio-visual failures. Experimental results show that commercial omni models such as Gemini substantially outperform other evaluators, yet even the strongest models continue to struggle with complex failures in our benchmark. These results demonstrate that MTAVG-Bench 2.0 provides a systematic benchmark for failure diagnosis in cinematic multi-talker audio-video generation.

音视频生成影视表现力故障诊断多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。