首次系统评估了不确定性量化在语义分割基础模型中的应用效果。
A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation

- 用SAM2编码器+轻量DPT解码器构建基线模型
- 四种不确定量化方法在多个数据集上表现各异
- 揭示精度、可靠性与计算开销间的权衡关系
基础模型正不断突破以往难以实现的精度与跨域泛化能力,但在可解释性、过度自信及对真实世界领域偏移的敏感性方面仍面临严峻挑战,尤其在安全与任务关键型应用中。不确定性量化(UQ)为应对这些问题提供了理论框架,但其在分割类基础模型中的集成尚未深入探索。本文首次系统评估了四种代表性UQ方法——蒙特卡洛丢弃、深度子集成、测试时增强和证据深度学习——应用于一个语义分割基础模型的表现。我们基于预训练的SAM2编码器微调了一个轻量级DPT解码器作为基准,并在Cityscapes、NYUv2以及两个具有挑战性的域外设置上进行测试。分析涵盖分割精度、校准度、不确定性质量与推理时间,揭示出预测性能、可靠性和计算成本之间的明显权衡。结果凸显了不确定性感知基础模型的潜力与当前局限,强调未来工作需协同优化准确性、鲁棒性与效率以支持实际部署。
原文摘要 · Abstract (English)
Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain generalization. Yet their lack of interpretability, tendency to be overconfident, and sensitivity to real-world domain shifts pose critical challenges for safety- and mission-critical applications. Uncertainty quantification (UQ) offers a principled way to address these issues, but its integration into segmentation foundation models has yet to be explored. In this paper we present the first systematic evaluation of UQ methods applied to a foundation model for semantic segmentation. We fine-tune a lightweight DPT decoder on top of the pretrained SAM2 encoder to establish a simple yet competitive baseline and benchmark four representative UQ approaches - Monte Carlo Dropout, Deep Sub-Ensemble, Test-Time Augmentation, and Evidential Deep Learning - across Cityscapes, NYUv2, and two challenging out-of-domain settings. Our analysis compares segmentation accuracy, calibration, uncertainty quality, and inference time, revealing clear trade-offs between predictive performance, reliability, and computational cost. These results highlight both the promise and the current limitations of uncertainty-aware foundation models, pointing to the need for future work that jointly optimizes accuracy, robustness, and efficiency for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。