arXiv:2506.07575cs.CVcs.LG2025-06被引 4

提出统一框架,让多模态大模型自知未知,提升推理可靠性。

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

  • 不依赖模型结构,通过提示扰动揭示多模态模型不确定性
  • 在18个基准上验证,可有效量化多模态响应的语义不确定性
  • 适合需要防幻觉、增强推理可信度的研究者与开发者

大型多模态模型(LMMs)虽融合多模态优势,被认为比纯语言大模型更鲁棒,但它们是否知道自己的未知?当前存在三大开放问题:如何统一评估不同LMM的不确定性,如何引导其暴露不确定性,以及如何为下游任务量化不确定性。为此,我们提出Uncertainty-o:(1) 一种模型无关框架,可揭示任意模态、架构、能力的LMM的不确定性;(2) 通过多模态提示扰动进行实证探索,揭示模型行为规律;(3) 推导出多模态语义不确定性公式,实现从多模态输出中量化不确定性。在涵盖多种模态的18个基准和10个LMM(含开源与闭源)上的实验表明,Uncertainty-o能可靠估计LMM不确定性,显著提升幻觉检测、幻觉缓解及不确定性感知的思维链推理等下游任务效果。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs), harnessing the complementarity among diverse modalities, are often considered more robust than pure Language Large Models (LLMs); yet do LMMs know what they do not know? There are three key open questions remaining: (1) how to evaluate the uncertainty of diverse LMMs in a unified manner, (2) how to prompt LMMs to show its uncertainty, and (3) how to quantify uncertainty for downstream tasks. In an attempt to address these challenges, we introduce Uncertainty-o: (1) a model-agnostic framework designed to reveal uncertainty in LMMs regardless of their modalities, architectures, or capabilities, (2) an empirical exploration of multimodal prompt perturbations to uncover LMM uncertainty, offering insights and findings, and (3) derive the formulation of multimodal semantic uncertainty, which enables quantifying uncertainty from multimodal responses. Experiments across 18 benchmarks spanning various modalities and 10 LMMs (both open- and closed-source) demonstrate the effectiveness of Uncertainty-o in reliably estimating LMM uncertainty, thereby enhancing downstream tasks such as hallucination detection, hallucination mitigation, and uncertainty-aware Chain-of-Thought reasoning.

多模态模型不确定性幻觉检测推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。