arXiv:2510.08581cs.SDcs.AI2025-10被引 1

测试语音查询下多模态模型的幻觉问题,发现噪声中错误率飙升30%。

Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions

  • 将图文基准转化为语音查询测试集,保持任务与标签一致
  • 语音输入使错误率上升3-6%(清晰)至30%(噪声环境)
  • 现有提示技术缓解有限,推动更可靠的语音交互系统研究

多模态模型中的幻觉问题已在图像-文本查询场景中被广泛研究,但语音查询对多模态幻觉的影响仍待探索,尽管语音接口日益普及。本文提出一种系统化流程,将现有多模态幻觉基准转换为语音查询版本,同时保留原任务和标签。我们在RePOPE上实现该流程,并发布RePOPE-Spk,其中所有查询均以不同声学条件下的语音音频形式呈现。实验结果表明,相比书面查询,语音输入显著加剧幻觉:在清晰语音下错误率上升3-6%,在环境噪声下最高提升30%。此外,多步提示与思维链推理仅能部分缓解该问题。研究结果为构建可靠语音交互系统及评估方法提供了新方向。

原文摘要 · Abstract (English)

Hallucinations in multimodal models have been extensively studied using benchmarks that probe reliability in image-text query settings. However, the effect of spoken queries on multimodal hallucinations remains largely unexplored, despite the growing role of voice interfaces. In this paper, we introduce a systematic pipeline that converts existing multimodal hallucination benchmarks into spoken-query versions while preserving the original tasks and labels. We instantiate this pipeline on RePOPE and release RePOPE-Spk, where all queries are provided as spoken audio under diverse input conditions. Experimental results show that hallucinations escalate when queries are spoken rather than written: error rates increase by 3-6% with clean speech and by up to 30% under environmental noise. Furthermore, many-shot prompting and chain-of-thought reasoning provide only partial mitigation. Our findings motivate new directions for building reliable voice interface systems and evaluations.

多模态语音查询幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。