系统研究分割任务中不确定性估计的影响因素,揭示其在不同场景下的有效性边界。
U-SEG: Uncertainty in SEGmentation -- A systematic multi-variable exploration

- 构建多变量实验框架,覆盖数据集、模型架构与下游任务
- 全景分割性能更差且泛化性不稳定,时间序列样本多数情况不划算
- 样本多样性仅在校准任务中有效,确定性模型对多数任务已足够
本研究深入探讨了不确定性估计与分割任务交叉领域中若干未充分研究的问题。已有研究表明,不确定性估计质量对多种变量高度敏感。由于不确定性估计常用于识别和处理实际应用中的预测错误,因此必须明确影响其表现的因素。例如:更具挑战性的任务或不同的数据集与模型架构是否会导致不确定性估计性能下降?视频序列中的前帧能否提供与其它方法相当的不确定性估计?能否结合多种不确定性估计方法,利用样本多样性获得更优结果?在何种情况下使用基于集成的方法比确定性网络更有优势?我们通过构建一个涵盖多种变量(如数据集、主干网络、下游任务)的大规模实验框架,对语义分割和全景分割进行了系统研究。结果表明:a)全景分割任务通常表现更差,且不同数据集与主干网络间性能方差大,说明泛化能力不可靠;b)时间序列样本在特定配置下可能有用,但多数情况下成本过高;c)样本多样性在下游校准任务中最具潜力,但在其他任务中难以超越简单替代方案;d)确定性方法对部分任务已足够,而集成方法在部署条件合适时可带来显著提升。
原文摘要 · Abstract (English)
In this study, we explore in depth a few under-studied topics at the intersection of uncertainty estimation and segmentation. Prior work has shown that the quality of uncertainty estimates can be very sensitive to a range of variables. As one of the main uses of uncertainty estimation is to help identify and deal with prediction errors in practical scenarios, any factors that affect this must be clearly identified. For example, do more challenging domains or different datasets and architectures result in worse performance when using uncertainty estimates? Can prior frames in a video sequence in fact provide useful uncertainty estimates comparable to other approaches? Is it possible to combine uncertainty estimation approaches, taking advantage of sample diversity, to get better estimates? Finally, when might it make sense to use an ensemble-based uncertainty estimate over a deterministic network? We address these questions by creating a framework for and executing a large scale study across many variables such as datasets, backbones, and downstream tasks, for both semantic and panoptic segmentation. We find that a) the more challenging task of panoptic segmentation usually results in worse performance while high performance variance between datasets and backbones indicates that generalization is not guaranteed, b) time series samples can be useful for specific configurations, but in many cases are not worth the cost, c) sample diversity shows the most promise in the downstream task of calibration, but otherwise fails to beat simpler alternatives, d) a deterministic approach is adequate for some downstream tasks, but ensembles allow for significant improvements if the right conditions can be achieved in deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。