构建多语言跨模态歧义消解基准,评估模型真实场景理解能力
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
- 设计双歧义数据集,图文互为线索实现唯一解释
- 19个主流模型表现远低于人类水平,差距显著
- 适合研究多模态推理与跨语言理解的学者使用
多模态大语言模型在视觉-语言任务中已取得显著进展,但在处理现实世界中语言与视觉语境的固有歧义时仍面临挑战。现有基准普遍忽略语言与视觉歧义,主要依赖单模态上下文消歧,未能发挥模态间相互澄清的潜力。为此,我们提出MUCAR,一个专为多语言跨模态歧义消解设计的新基准。MUCAR包含两个部分:一是多语言数据集,其中模糊文本表达由对应视觉内容唯一解析;二是双歧义数据集,系统性地将模糊图像与模糊文本配对,每组通过模态间互推达成单一明确解释。对19种先进多模态模型(含开源与闭源架构)的广泛评估显示,其性能与人类水平存在显著差距,凸显未来需发展更复杂的跨模态歧义理解方法,推动多模态推理边界进一步拓展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text pairs with clear and explicit meanings. However, resolving the inherent ambiguities present in real-world language and visual contexts remains a challenge. Existing multimodal benchmarks typically overlook linguistic and visual ambiguities, relying mainly on unimodal context for disambiguation and thus failing to exploit the mutual clarification potential between modalities. To bridge this gap, we introduce MUCAR, a novel and challenging benchmark designed explicitly for evaluating multimodal ambiguity resolution across multilingual and cross-modal scenarios. MUCAR includes first a multilingual dataset where ambiguous textual expressions are uniquely resolved by corresponding visual contexts, and second a dual-ambiguity dataset that systematically pairs ambiguous images with ambiguous textual contexts, with each combination carefully constructed to yield a single, clear interpretation through mutual disambiguation. Extensive evaluations involving 19 state-of-the-art multimodal models--encompassing both open-source and proprietary architectures--reveal substantial gaps compared to human-level performance, highlighting the need for future research into more sophisticated cross-modal ambiguity comprehension methods, further pushing the boundaries of multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。