arXiv:2603.26349stat.MLcs.AI2026-03

用生成模型估算多模态数据不确定性,提升预测可信度。

Generative Score Inference for Multimodal Data

  • 用生成模型合成样本逼近条件得分分布,灵活量化不确定性。
  • 在大模型幻觉检测和图像描述不确定性估计中表现领先。
  • 适合需要高可信度决策的多模态应用,如医疗影像分析。

准确的不确定性量化对各类监督学习任务至关重要,尤其在图像、文本等复杂多模态数据场景下。现有方法常受限于严格假设和泛化能力不足。为此,我们提出生成得分推断(Generative Score Inference, GSI),一种灵活的推断框架,可在多种多模态学习问题中构建统计有效且信息丰富的预测集与置信集。GSI利用深度生成模型生成的合成样本近似条件得分分布,实现无需强假设的精确不确定性量化。我们在两个典型场景中验证其性能:大语言模型的幻觉检测与图像描述的不确定性估计。结果表明,GSI在幻觉检测中达到当前最优表现,并在图像描述中实现稳健的预测不确定性估计;其性能随底层生成模型质量提升而增强。这些发现凸显了GSI作为通用推断框架的潜力,显著提升了多模态学习中的不确定性量化与可信度。

原文摘要 · Abstract (English)

Accurate uncertainty quantification is crucial for making reliable decisions in various supervised learning scenarios, particularly when dealing with complex, multimodal data such as images and text. Current approaches often face notable limitations, including rigid assumptions and limited generalizability, constraining their effectiveness across diverse supervised learning tasks. To overcome these limitations, we introduce Generative Score Inference (GSI), a flexible inference framework capable of constructing statistically valid and informative prediction and confidence sets across a wide range of multimodal learning problems. GSI utilizes synthetic samples generated by deep generative models to approximate conditional score distributions, facilitating precise uncertainty quantification without imposing restrictive assumptions about the data or tasks. We empirically validate GSI's capabilities through two representative scenarios: hallucination detection in large language models and uncertainty estimation in image captioning. Our method achieves state-of-the-art performance in hallucination detection and robust predictive uncertainty in image captioning, and its performance is positively influenced by the quality of the underlying generative model. These findings underscore the potential of GSI as a versatile inference framework, significantly enhancing uncertainty quantification and trustworthiness in multimodal learning.

不确定性量化多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。