arXiv:2608.27135cs.CL2026-08被引 1

研究语音与文本输入对多模态模型判断一致性的影响

Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models

论文配图:Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models
图 1 · 摘自论文原文
  • 构建跨模态对比三元组基准,覆盖18个MENA国家10,150张文化图像
  • 语音输入使模型在三元组任务中错误率上升,部分失败更明显
  • 发现语言和模态切换引发显著不一致,适合关注AI公平性的研究者

多模态基础模型越来越多地用于以语音为主的助手,需理解口语查询并生成视觉相关的决策。然而,语义等价的查询在不同模态(文本与语音)和语言(英语与阿拉伯语)下是否产生一致判断仍不清楚。本文引入一个语音增强的视觉定位对比三元组基准,涵盖来自18个MENA国家的10,150张文化相关图像,每张图像配有一条支持陈述和两条看似合理但不支持的备选陈述。将对比不稳定性定义为模型无法解决三元组内所有陈述的条件概率,从而分离出碎片化推理与完全失败。在英语和阿拉伯语的文本与语音输入下评估多个近期多模态模型,发现模态与语言切换引入了显著的三元组级不一致,且这些不一致无法通过整体准确率充分捕捉,语音会放大部分失败。该基准已公开提供给社区使用。

原文摘要 · Abstract (English)

Multimodal foundation models are increasingly used in speech-first assistants that must interpret spoken queries and produce visually grounded decisions. Yet it remains unclear whether semantically equivalent queries yield consistent judgments across modality (text vs. speech) and language (English vs. Arabic). We introduce a speech-augmented visually grounded contrastive triplet benchmark spanning 10,150 culturally grounded images from 18 MENA countries, where each image is paired with one supported statement and two plausible but unsupported alternatives. We define contrastive instability as the conditional rate at which a model fails to resolve all statements within a triplet, isolating fragmented reasoning from complete failure. Evaluating recent multimodal models under text and speech in English and Arabic, we find that modality and language shifts introduce substantial triplet-level inconsistencies that are not fully captured by aggregate accuracy, with speech amplifying partial failures. We make the benchmark publicly available to the community.

多模态语音识别模型不一致公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。