用统计方法让视觉语言模型生成内容更可靠,大幅减少幻觉。
Towards Statistical Factuality Guarantee for Large Vision-Language Models

- 将模型输出视为假设,通过统计检验筛选可信信息
- 在场景描述中将错误率从87.8%降至10.0%,真阳性率达95.3%
- 适用于任意黑箱模型和任务,提供可证明的幻觉控制保障
大型视觉语言模型(LVLM)在图像引导的自由文本生成任务中表现优异,但其生成内容与视觉上下文不符的幻觉问题日益严重,阻碍了高可靠性应用。本文提出ConfLVLM框架,基于合流预测实现有限样本、分布无关的统计事实性保证。该框架将LVLM视为假设生成器,每个生成的文本细节作为独立假设,利用高效的启发式不确定性度量进行统计检验,过滤不可靠陈述后再返回结果。在通用场景理解、医学影像报告生成和文档理解三个典型领域进行实验,结果显示,对LLaVa-1.5的场景描述,错报率从87.8%降至10.0%,真阳性率为95.3%。结果表明,ConfLVLM具有高度灵活性,可适配任意黑箱LVLM与任意不确定性度量,在任意图像条件下的自由文本生成任务中提供严格的幻觉风险控制。
原文摘要 · Abstract (English)
Advancements in Large Vision-Language Models (LVLMs) have demonstrated promising performance in a variety of vision-language tasks involving image-conditioned free-form text generation. However, growing concerns about hallucinations in LVLMs, where the generated text is inconsistent with the visual context, are becoming a major impediment to deploying these models in applications that demand guaranteed reliability. In this paper, we introduce a framework to address this challenge, ConfLVLM, which is grounded on conformal prediction to achieve finite-sample distribution-free statistical guarantees on the factuality of LVLM output. This framework treats an LVLM as a hypothesis generator, where each generated text detail (or claim) is considered an individual hypothesis. It then applies a statistical hypothesis testing procedure to verify each claim using efficient heuristic uncertainty measures to filter out unreliable claims before returning any responses to users. We conduct extensive experiments covering three representative application domains, including general scene understanding, medical radiology report generation, and document understanding. Remarkably, ConfLVLM reduces the error rate of claims generated by LLaVa-1.5 for scene descriptions from 87.8\% to 10.0\% by filtering out erroneous claims with a 95.3\% true positive rate. Our results further demonstrate that ConfLVLM is highly flexible, and can be applied to any black-box LVLMs paired with any uncertainty measure for any image-conditioned free-form text generation task while providing a rigorous guarantee on controlling the risk of hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。