arXiv:2608.29193cs.LG2026-08中稿 · EMNLP

提出HalluPrism诊断框架,区分模型失败类型与是否该拒绝回答。

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

论文配图:HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide
图 1 · 摘自论文原文
  • 通过图像退化、空白替换和语义检查重运行答案,识别失败模式。
  • 图像移除后仍高自信的现象最常见,语义不稳定性更利于区分错误类型。
  • 诊断签名可提升故障分类效果,但不自动提高答案正确性排序。

多模态大语言模型对不同原因导致的错误可能赋予相似置信度。本文提出HalluPrism,一种行为诊断方法:在图像退化、空白图替换及语义/关系检查后重新运行答案。通过三类探针生成视觉扰动敏感性(V)、图像移除后置信保留率(L)和语义/关系探针不稳定性(A)的联合签名。在四个基准、四款MLLM上共58,000+样本测试中,图像移除后置信保留最普遍;语义不稳定性更能区分错误类别。48组源-目标检查中仅18组对角线对齐,说明坐标需联合解读而非独立因果。固定数据集下,联合签名使HallusionBench的故障家族AUROC从0.634升至0.769,VizWiz从0.707升至0.817,POPE与VSR提升较小。联合XGBoost分析显示,从单标量置信度的0.78提升至使用(V,L,A)的0.95,加入置信度后达0.97。同一签名无法自动改善答案正确性排序,三种直接标量化方式甚至可能损害它。结果表明,应先诊断失败结构,再决定是否放弃或修正。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confidence retention (L), and grounding/relation-probe instability (A). Across 58K+ examples from four benchmarks and four MLLMs, image-removal confidence retention is most prevalent, while grounding/relation-probe instability better separates failure families. Only 18 of 48 source-target checks are diagonally aligned, so the coordinates should be interpreted jointly rather than as independent causal sources. With the dataset fixed, the joint signature improves failure-family AUROC from 0.634 to 0.769 on HallusionBench and from 0.707 to 0.817 on VizWiz, with smaller gains on POPE and VSR. In pooled XGBoost analysis, AUROC rises from 0.78 with scalar confidence to 0.95 with (V, L, A) and 0.97 when confidence is added. The same signature does not automatically improve correctness ranking. The three tested direct scalarizations can harm it. These results separate failure diagnosis from abstention scoring: multimodal uncertainty should characterize failure structure before it is used to decide whether to abstain or correct.

多模态模型诊断不确定性幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。