arXiv:2509.16648cs.AIcs.CL2025-09EMNLP被引 5

通过等效与互补采样,评估多模态大模型预测可信度。

FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs

  • 设计等效与互补输入采样,探测模型一致性与敏感性。
  • 在视觉和音频任务上,选择性预测性能提升33.3%和29.6%。
  • 无需真值标签,仅需黑盒访问,适用于各类现成多模态模型。

多模态大语言模型(MLLMs)的可信度评估因多模态输入形式多样而具有挑战性。本文提出功能等效采样信任评估方法(FESTA),一种基于等效与互补输入采样的多模态输入采样技术,用于生成不确定性度量。该任务保持型采样方法扩展输入空间,以探测模型的一致性(通过等效样本)和敏感性(通过互补样本)。FESTA仅需模型的输入输出访问(黑盒),无需真实标签(无监督)。在多种现成多模态大模型上,针对视觉和音频推理任务的实验表明,提出的不确定性估计在检测误预测方面显著提升选择性预测性能:视觉-大模型相对提升33.3%,音频-大模型相对提升29.6%(以受试者工作特征曲线下面积,AUROC为指标)。代码已开源。

原文摘要 · Abstract (English)

The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging due to the diverse multi-modal input paradigms. We propose Functionally Equivalent Sampling for Trust Assessment (FESTA), a multimodal input sampling technique for MLLMs, that generates an uncertainty measure based on the equivalent and complementary input samplings. The proposed task-preserving sampling approach for uncertainty quantification expands the input space to probe the consistency (through equivalent samples) and sensitivity (through complementary samples) of the model. FESTA uses only input-output access of the model (black-box), and does not require ground truth (unsupervised). The experiments are conducted with various off-the-shelf multi-modal LLMs, on both visual and audio reasoning tasks. The proposed FESTA uncertainty estimate achieves significant improvement (33.3% relative improvement for vision-LLMs and 29.6% relative improvement for audio-LLMs) in selective prediction performance, based on area-under-receiver-operating-characteristic curve (AUROC) metric in detecting mispredictions. The code implementation is open-sourced.

多模态可信度评估不确定性量化黑盒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。