通过专家标注差异提升显微图像目标检测器的置信度校准
Leveraging Multi-Rater Annotations to Calibrate Object Detectors in Microscopy Imaging
- 分别训练各专家标注数据上的模型,集成预测模拟共识
- 在结直肠类器官数据集上校准性能提升,检测精度不变
- 适合需要可信置信度的生物医学图像分析场景
基于深度学习的目标检测器在显微成像中表现优异,但其置信度估计常缺乏校准,限制了在生物医学应用中的可靠性。本文提出一种新方法,利用多专家标注提升模型校准能力:分别在单个专家的标注数据上训练模型,并集成其预测以模拟专家共识。相比混合标注的标签采样策略,该方法更系统地捕捉了专家间差异。在两位专家标注的结直肠类器官数据集上实验表明,该专家特定集成策略显著改善了校准性能,同时保持了相当的检测准确率。结果表明,显式建模标注者分歧可提升生物医学图像中目标检测器的可信度。
原文摘要 · Abstract (English)
Deep learning-based object detectors have achieved impressive performance in microscopy imaging, yet their confidence estimates often lack calibration, limiting their reliability for biomedical applications. In this work, we introduce a new approach to improve model calibration by leveraging multi-rater annotations. We propose to train separate models on the annotations from single experts and aggregate their predictions to emulate consensus. This improves upon label sampling strategies, where models are trained on mixed annotations, and offers a more principled way to capture inter-rater variability. Experiments on a colorectal organoid dataset annotated by two experts demonstrate that our rater-specific ensemble strategy improves calibration performance while maintaining comparable detection accuracy. These findings suggest that explicitly modelling rater disagreement can lead to more trustworthy object detectors in biomedical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。