用统计方法让大模型回答更靠谱,避免胡说八道。
Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction
- 基于分割共形预测框架,动态调整置信度阈值。
- 在科学问答和多模态评测中,错误率始终低于设定风险水平α。
- 无需重训练,适合医疗、自动驾驶等高安全场景使用。
本研究针对大型视觉语言模型(LVLMs)在视觉问答任务中的幻觉问题,提出一种基于分割共形预测(SCP)的框架。尽管LVLMs在多模态推理上表现优异,但其输出常包含高置信度的幻觉内容,危及安全关键应用。本文提出一种模型无关的不确定性量化方法,结合动态阈值校准与跨模态一致性验证。通过将数据划分为校准集和测试集,计算非符合度得分,构建具有统计保证的预测集,满足用户定义的风险水平(α)。核心创新包括:(1) 严格控制边缘覆盖性,确保经验误差率始终低于α;(2) 预测集大小随α动态反向调整,过滤低置信度输出;(3) 无需先验分布假设或重新训练。在ScienceQA、MMMU等基准上,对八种LVLM进行评估,结果表明SCP在所有α值下均满足理论保证。该框架在不同校准-测试划分比例下表现稳定,展现出在医疗、自动驾驶等安全敏感领域部署的鲁棒性。本工作弥合了多模态AI系统中理论可靠性与实际可用性之间的鸿沟,提供了一种可扩展的幻觉检测与不确定性感知决策方案。
原文摘要 · Abstract (English)
This study addresses the critical challenge of hallucination mitigation in Large Vision-Language Models (LVLMs) for Visual Question Answering (VQA) tasks through a Split Conformal Prediction (SCP) framework. While LVLMs excel in multi-modal reasoning, their outputs often exhibit hallucinated content with high confidence, posing risks in safety-critical applications. We propose a model-agnostic uncertainty quantification method that integrates dynamic threshold calibration and cross-modal consistency verification. By partitioning data into calibration and test sets, the framework computes nonconformity scores to construct prediction sets with statistical guarantees under user-defined risk levels ($α$). Key innovations include: (1) rigorous control of \textbf{marginal coverage} to ensure empirical error rates remain strictly below $α$; (2) dynamic adjustment of prediction set sizes inversely with $α$, filtering low-confidence outputs; (3) elimination of prior distribution assumptions and retraining requirements. Evaluations on benchmarks (ScienceQA, MMMU) with eight LVLMs demonstrate that SCP enforces theoretical guarantees across all $α$ values. The framework achieves stable performance across varying calibration-to-test split ratios, underscoring its robustness for real-world deployment in healthcare, autonomous systems, and other safety-sensitive domains. This work bridges the gap between theoretical reliability and practical applicability in multi-modal AI systems, offering a scalable solution for hallucination detection and uncertainty-aware decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。