提出无需训练的多模型不确定性量化框架,提升多智能体推理可靠性。
CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

- 将多个视觉语言模型输出映射到统一语义空间,聚合为群体意见
- 集体不确定性和模型间冲突(JSD)可分别衡量系统级置信度与分歧
- 适用于开源与商用模型,在小模型和商业模型场景均表现优异
集成异构视觉语言模型(VLMs)能增强多模态推理,但个体模型置信度或聚合答案无法反映系统级可靠性。我们提出CUSP(通过语义意见池化实现集体不确定性),一种无需训练的不确定性量化框架:将多个VLM响应映射至共享语义空间,聚合为群体语义意见,并报告两个互补的系统级信号——集体不确定性(群体意见离散度)和詹森-香农散度(JSD,模型间意见冲突)。在该语义空间中,未归一化的集体熵恰好分解为各模型语义熵的均值与JSD,分离了总体离散与模型冲突。无需令牌对数或校准标签,CUSP可应用于开放权重与商业VLM。静态多模型集成中,集体不确定性在小模型阶段表现最优(预测错误检测AUROC达0.764,弃权效果AUARC达0.889),优于多数投票与朴素选择4.7至15.8点,且随集成规模扩大优势更明显;JSD在商业模型阶段最强(AUROC 0.819,AUARC 0.910),可有效识别高难度问题(最高AUROC达0.982)。聚合预测性能也优于单个模型5.6至13.0点。在多步多智能体系统全程中,子智能体集体不确定性在系统故障预警上显著优于随机水平(AUROC 0.619),且在弃权排序中表现最佳(AUARC 0.699)。
原文摘要 · Abstract (English)
Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level opinions. Within this pooled semantic opinion, the unnormalized collective entropy decomposes exactly into the mean of the models' individual semantic entropies and the JSD, separating total dispersion from model conflict. Requiring neither token logits nor calibration labels, CUSP applies to open-weight and commercial VLMs alike. In static multi-VLM ensembles, collective uncertainty is the strongest signal in the small-model regime (0.764 AUROC for prediction-error detection, 0.889 AUARC for abstention), outperforming uncertainty baselines majority voting and naive selection by 4.7 to 15.8 points and widening its margin as the ensemble grows; JSD is strongest in the evaluated commercial regime (0.819 AUROC, 0.910 AUARC) and ranks hard-answer model conflict with AUROC up to 0.982. The pooled prediction also improves accuracy over the average single model by 5.6 to 13.0 points. Over the full trajectory of a multi-step, multi-agent system, subagent collective uncertainty ranks system failures above chance (0.619 AUROC) and gives the best abstention ordering among the evaluated signals (0.699 AUARC).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。