联邦学习在牙科影像分割中抗噪声与标注错误,且可检测异常数据源。
Impact of Labeling Inaccuracy and Image Noise on Tooth Segmentation in Panoramic Radiographs using Federated, Centralized and Local Learning
- 通过客户端损失曲线监控,实现对异常数据源的自动检测。
- 在标注缺失、图像噪声等干扰下,联邦学习仍保持94.8%以上的分割精度。
- 适合需要保护隐私的多机构医疗AI系统部署,尤其适用于数据质量不一场景。
目的:联邦学习(FL)可缓解隐私限制、数据质量异质性及标注不一致问题,在牙科诊断AI中具有潜力。本文对比了联邦学习(FL)、集中式学习(CL)与本地学习(LL)在全景牙片中牙齿分割的表现,涵盖四种数据污染场景:基准(未修改数据)、标注扰动(扩张/缺失标注)、图像质量扰动(添加高斯噪声),以及排除一个存在缺陷的数据客户端。采用注意力U-Net模型,在六家机构共2066张放射影像上训练,使用Flower框架实现联邦学习。监测各客户端的训练与验证损失轨迹以进行异常检测,并在独立测试集上评估Dice、IoU、HD、HD95和ASSD等指标。通过威尔科克斯符号秩检验分析显著性。结果:基准条件下,FL中位Dice为0.94889(ASSD: 1.33229),略优于CL的0.94706(ASSD: 1.37074),显著优于LL的0.93557–0.94026(ASSD: 1.51910–1.69777)。标注扰动下,FL维持最高中位数Dice 0.94884(ASSD: 1.46487),优于CL的0.94183(ASSD: 1.75738)和LL的0.93003–0.94026(ASSD: 1.51910–2.11462)。图像噪声下,FL表现最佳,Dice为0.94853(ASSD: 1.31088),CL为0.94787(ASSD: 1.36131),LL为0.93179–0.94026(ASSD: 1.51910–1.77350)。排除故障客户端后,FL Dice达0.94790(ASSD: 1.33113),优于CL的0.94550(ASSD: 1.39318)。损失曲线监控可有效识别异常站点。结论:联邦学习在多种数据污染场景下表现匹配或优于集中式学习,显著优于本地学习,同时保障隐私。客户端损失轨迹提供有效的异常检测机制,支持其作为可扩展临床AI部署的实用方案。
原文摘要 · Abstract (English)
Objectives: Federated learning (FL) may mitigate privacy constraints, heterogeneous data quality, and inconsistent labeling in dental diagnostic AI. We compared FL with centralized (CL) and local learning (LL) for tooth segmentation in panoramic radiographs across multiple data corruption scenarios. Methods: An Attention U-Net was trained on 2066 radiographs from six institutions across four settings: baseline (unaltered data); label manipulation (dilated/missing annotations); image-quality manipulation (additive Gaussian noise); and exclusion of a faulty client with corrupted data. FL was implemented via the Flower AI framework. Per-client training- and validation-loss trajectories were monitored for anomaly detection and a set of metrics (Dice, IoU, HD, HD95 and ASSD) was evaluated on a hold-out test set. From these metrics significance results were reported through Wilcoxon signed-rank test. CL and LL served as comparators. Results: Baseline: FL achieved a median Dice of 0.94889 (ASSD: 1.33229), slightly better than CL at 0.94706 (ASSD: 1.37074) and LL at 0.93557-0.94026 (ASSD: 1.51910-1.69777). Label manipulation: FL maintained the best median Dice score at 0.94884 (ASSD: 1.46487) versus CL's 0.94183 (ASSD: 1.75738) and LL's 0.93003-0.94026 (ASSD: 1.51910-2.11462). Image noise: FL led with Dice at 0.94853 (ASSD: 1.31088); CL scored 0.94787 (ASSD: 1.36131); LL ranged from 0.93179-0.94026 (ASSD: 1.51910-1.77350). Faulty-client exclusion: FL reached Dice at 0.94790 (ASSD: 1.33113) better than CL's 0.94550 (ASSD: 1.39318). Loss-curve monitoring reliably flagged the corrupted site. Conclusions: FL matches or exceeds CL and outperforms LL across corruption scenarios while preserving privacy. Per-client loss trajectories provide an effective anomaly-detection mechanism and support FL as a practical, privacy-preserving approach for scalable clinical AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。