无需训练和外部工具,用模型内部特征量化多模态大模型输出不确定性。
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
- 基于模型自身特征计算语义多样性与响应不一致性,融合全局与局部信息。
- 在图像、音频、视频等多任务中均优于基线方法,尤其在对抗样本和分布外数据上表现突出。
- 适用于文本、图像、音频生成等多种输出任务,适合部署于高可靠性要求场景。
尽管多模态大语言模型(MLLMs)具备强大能力,但可能产生看似合理却错误的输出,阻碍其可靠应用。准确的不确定性度量可帮助将不可靠查询转交人工专家或更大模型以提升性能。然而,现有度量方法存在局限:仅适用于特定模态、依赖外部工具或计算开销大。本文提出UMPIRE,一种无需训练、不依赖外部工具的多模态大模型不确定性量化框架,仅利用模型内部模态特征,在多种输入输出模态下高效运行。UMPIRE通过计算给定任务实例下采样响应的不一致调整语义体积,有效捕捉样本的全局语义多样性与局部响应不一致性,基于模型内部置信度。我们提出了多模态大模型的不确定性理想标准,并提供理论分析支持设计。大量实验表明,UMPIRE在图像、音频、视频-文本基准测试中持续优于基线指标,在对抗性与分布外设置下表现优异。此外,我们验证了UMPIRE在非文本输出任务(如图像和音频生成)中的泛化能力。
原文摘要 · Abstract (English)
Despite their capabilities, Multimodal Large Language Models (MLLMs) may produce plausible but erroneous outputs, hindering reliable deployment. Accurate uncertainty metrics could enable escalation of unreliable queries to human experts or larger models for improved performance. However, existing uncertainty metrics have practical constraints, such as being designed only for specific modalities, reliant on external tools, or computationally expensive. We introduce UMPIRE, a training-free uncertainty quantification framework for MLLMs that works efficiently across various input and output modalities without external tools, relying only on the models' own internal modality features. UMPIRE computes the incoherence-adjusted semantic volume of sampled MLLM responses for a given task instance, effectively capturing both the global semantic diversity of samples and the local incoherence of responses based on internal model confidence. We propose uncertainty desiderata for MLLMs and provide theoretical analysis motivating UMPIRE's design. Extensive experiments show that UMPIRE consistently outperforms baseline metrics in error detection and uncertainty calibration across image, audio, and video-text benchmarks, including adversarial and out-of-distribution settings. We also demonstrate UMPIRE's generalization to non-text output tasks, including image and audio generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。