arXiv:2603.23938cs.CL2026-03被引 1

评测多模态模型能否根据上下文准确发声,发现现有模型在语音生成上仍有短板。

OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

  • 设计新基准,要求模型根据语音指令、文本和图像朗读出符合语境的语音。
  • 测试8个模型均表现不佳,尤其在情感、口音等6类声学特征控制上薄弱。
  • 揭示模型语音生成失败三大原因,助力开发能有效‘开口说话’的多模态系统。

当前多数多模态模型评测依赖文本输出,难以判断其是否真正具备语音表达能力。为此,我们提出OmniACBench,一个用于评估多模态模型在上下文约束下声学控制能力的基准。给定一段语音指令、文本脚本和一张图像,模型需以恰当的语调与风格朗读脚本。OmniACBench包含3,559个经验证的实例,覆盖六种声学特征:语速、发声方式、发音、情感、整体口音和音色。对八个模型的实验表明,尽管它们在传统文本评测中表现优异,但在该设定下仍存在明显不足。分析显示,主要瓶颈并非单模态处理,而在于多模态上下文融合以生成忠实语音。此外,我们识别出三种常见失效模式:控制力弱、隐含推理失败、多模态对齐失败,为构建能有效‘口语化回答’的模型提供了关键洞察。

原文摘要 · Abstract (English)

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for evaluating context-grounded acoustic control in omni-modal models. Given a spoken instruction, a text script, and an image, a model must read the script aloud with an appropriate tone and manner. OmniACBench comprises 3,559 verified instances covering six acoustic features: speech rate, phonation, pronunciation, emotion, global accent, and timbre. Extensive experiments on eight models reveal their limitations in the proposed setting, despite their strong performance on prior textual-output evaluations. Our analyses show that the main bottleneck lies not in processing individual modalities, but in integrating multimodal context for faithful speech generation. Moreover, we identify three common failure modes-weak direct control, failed implicit inference, and failed multimodal grounding-providing insights for developing models that can verbalize responses effectively.

多模态语音生成评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。