提出新框架解决3D医疗图像分割中的过度自信问题。
Are We Overconfident in Models and Results for Semi-Supervised 3D Medical Image Segmentation?

- 分离置信度与不确定性,纠正预测偏差
- 在三个数据集上保持稳定性能,避免过拟合
- 建议多轮测试报告结果,提升评估可靠性
半监督学习已成为降低标注成本的主流方法。然而,我们指出当前进展被双重过度自信问题所掩盖:算法层面,主流伪标签框架常将预测置信度误认为不确定性,导致严重确认偏差;策略层面,因多个基准数据集缺乏独立验证集,部分研究使用测试集进行验证,造成性能估计虚高。后续方法为超越报告结果,被迫采用相同策略,引发过拟合竞赛。这令人担忧,领域内看似显著的数值提升可能源于过拟合而非真实进步。为此,我们提出基于双轴可靠性评估机制的三空间校准分割框架,显式解耦置信度与不确定性,在特征、概率和图像空间协同检测并修正确认偏差。在三个基准数据集上,TCSeg 在现有评估协议下持续表现强劲。更重要的是,我们倡导社区采用多轮运行协议报告最终检查点结果,以建立更严格、更真实的基准。代码将公开:github.com/DirkLiii/TCSeg。
原文摘要 · Abstract (English)
Semi-supervised learning has become a dominant paradigm for reducing annotation costs. However, we argue that the current progress is clouded by a twofold overconfidence problem. Algorithmically, mainstream pseudo-labeling frameworks often conflate prediction confidence with uncertainty, leading to severe confirmation bias. Strategically, since multiple benchmark datasets lack dedicated validation sets, some studies use the test set for validation as well, leading to inflated performance estimates. Subsequent methods, compelled to employ the same strategy to surpass reported SOTA, trigger an arms race of overfitting. This raises concerns that the impressive numerical gains in the community may reflect overfitting rather than genuine progress. Thus, we propose a tri-space calibrated segmentation framework founded on a principled dual-axis reliability assessment engine. It explicitly decouples confidence from uncertainty and uses this signal to detect and correct confirmation bias across feature, probability, and image spaces in a collaborative manner. Across three benchmark datasets, TCSeg consistently delivers strong performance under existing evaluation protocols. More importantly, we advocate that the community report final-checkpoint results under multiple-run protocols, thereby establishing more rigorous benchmarks with a more realistic perspective. Code will be available: github.com/DirkLiii/TCSeg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。