让多模态大模型学会识别并应对不确定,提升关键决策可靠性。
Uncertainty-Aware Decision Making in Multimodal Large Language Models

- 从感知、语义、推理等多源出发,构建决策导向的不确定性评估框架。
- 提出用不确定性信号指导模型主动澄清、拒绝回答或调用外部验证。
- 适合关注多模态模型安全、可信与鲁棒性的研究者与开发者。
多模态大语言模型(MLLMs)越来越多地回答依赖视觉、文本、时间、音频、文档、图表或具身证据的问题。其错误不仅源于语言能力不足,更可能来自输入质量差、感知偏差、语义漂移、模态冲突、推理不稳、分布偏移或问题本身不可答。本文围绕以决策为中心的框架,系统梳理不确定性感知的文献:不确定性来源产生可观察信号,信号需校准以控制风险,校准后的不确定性应决定系统行为。涵盖词汇与对数概率不确定性、语义分歧、扰动不稳定性、定位与归因分数、口语化置信度、验证器与评判分数、共形预测、选择性回答、拒答、澄清、检索、自检与升级等方法。核心观点是:不确定性不应仅作为置信度数值评估,而应看其是否在证据不足、冲突、分布偏移或高风险情境下改善系统行为。本文区别于纯文本不确定性、泛化性综述、幻觉研究及安全导向回顾,并指出未来挑战:源意识分解、行动导向基准、分布偏移下的校准、黑盒不确定性估计、更多模态覆盖、可复现报告及以人为中心的不确定性传达。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a question that is not answerable from the supplied evidence. This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action. We review work on token and logit uncertainty, semantic disagreement, perturbation instability, grounding and attribution scores, verbalized confidence, verifier and judge scores, conformal prediction, selective answering, abstention, clarification, retrieval, self-checking, and escalation. The central argument is that uncertainty should not be evaluated only as a confidence number; it should be evaluated by whether it improves behavior under insufficient, conflicting, shifted, or high-risk multimodal evidence. We position this survey against text-only uncertainty and abstention surveys, broad MLLM surveys, MLLM hallucination surveys, and safety-oriented reviews. We conclude with open problems in source-aware decomposition, action-aware benchmarks, calibration under shift, black-box uncertainty estimation, broader modality coverage, reproducible reporting, and human-centered uncertainty communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。