arXiv:2602.01561cs.CVcs.AI2026-02

构建多模态反常识数据集,测试模型在非常规场景下的推理能力。

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

  • 通过检索增强的上下文学习框架,让小模型复用大模型的推理能力。
  • 在非常规场景下相比基线提升8.3%,验证了方法有效性。
  • 适合研究视觉语言模型鲁棒性与跨文化适应性的学者使用。

多模态常识推理仍是人工智能中的基础挑战。我们提出多模态反常识基准(Multimodal UNcommonsense, MUN),用于评估模型在偏离典型视觉或语境预期情境下的表现。MUN将视觉场景与自然语言描述的意外或不可能结果配对,促使模型要么用日常逻辑解释看似奇怪的画面,要么从普通场景中发现出人意料的解读。为此,我们设计了一种基于检索的上下文学习(R-ICL)框架,无需额外训练即可将大模型的推理能力迁移到小模型。利用新颖的多模态集成检索器(MER),该方法能在图像与文本对被刻意错配时仍识别出语义相关的示例。实验表明,该方法在平均上比基线ICL方法提升8.3%,凸显了R-ICL在低频、非典型场景中的有效性。MUN为评估和提升视觉语言模型在真实世界、文化多样性及非典型场景下的鲁棒性与适应性开辟了新方向。

原文摘要 · Abstract (English)

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios that deviate from typical visual or contextual expectations. MUN pairs visual scenes with surprising or unlikely outcomes described in natural language, prompting models to either rationalize seemingly odd images using everyday logic or uncover unexpected interpretations in ordinary scenes. To support this task, we propose a retrieval-based in-context learning (R-ICL) framework that transfers reasoning capabilities from larger models to smaller ones without additional training. Leveraging a novel Multimodal Ensemble Retriever (MER), our method identifies semantically relevant exemplars even when image and text pairs are deliberately discordant. Experiments show an average improvement of 8.3% over baseline ICL methods, highlighting the effectiveness of R-ICL in low-frequency, atypical settings. MUN opens new directions for evaluating and improving visual-language models' robustness and adaptability in real-world, culturally diverse, and non-prototypical scenarios.

多模态常识推理检索增强鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。