通过分析视觉证据来源与分布,提升多模态模型闭合答案的可信度检测。
Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers

- 逐层分析支持或反驳预测的视觉贡献,构建可解释的证据路径。
- 在4个基准上平均错误准确率提升0.134,所有12组实验均获增益。
- 无需扰动图像、反向传播或外部验证器,适用于白盒闭合系统。
多模态大语言模型(MLLM)的可靠部署需要判断一个自信的视觉回答是否可信、需审核或应转至更强系统。置信度分数虽能反映候选答案的边际,却无法揭示这些边际所依赖的视觉读数来源及其分布。本文研究基于生成答案时相同白盒预填充路径的推理阶段风险检测。见证证据组合(WEP)逐层估计哪些视觉贡献支持或反驳预测候选项,并通过两类可解释路径汇总:与问题相关的证据溯源和有符号证据集中度。嵌套分组验证选择更可靠的路径家族及稀疏的top-k路径组合,再与候选置信度融合。WEP无需图像扰动、解码改动、反向传播或外部验证器。在三个MLLM和四个二分类答案基准上,平均错误准确率提升0.134。全部12组模型-数据集组合均有提升,10对中图像聚类自助区间均为正。该方法针对白盒闭合答案系统,使用带标签校准片段。
原文摘要 · Abstract (English)
Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Confidence scores capture candidate margins, but not where the estimated signed visual readouts associated with those margins come from or how they are distributed. We study inference-time risk detection for closed visual answers using the same white-box prefill path that produces the answer. Witness Evidence Portfolios (WEP) first estimates, layer by layer, which visual contributions support or contradict the predicted candidate. It summarizes these contributions through two interpretable route families: question-related evidence provenance and signed evidence concentration. Nested grouped validation chooses the more reliable family and a sparse top-k route portfolio, which is fused with candidate confidence. WEP needs no image perturbation, decoding change, backward pass, or external verifier. Across three MLLMs and four binary-answer benchmarks, WEP improves mean error AP by 0.134. All 12 model--dataset gains are positive, and image-cluster bootstrap intervals are strictly positive on 10 pairs. WEP targets white-box closed-answer systems and uses a labeled calibration slice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。