用解剖知识检测生成图像中的异常形状,提升合成数据可信度。
Knowledge-based anomaly detection for identifying network-induced shape artifacts
- 基于角度梯度分布构建解剖边界特征空间,再用孤立森林检测异常
- 在两个乳腺影像合成数据集上AUC达0.97和0.91,精准定位异常区域
- 结果与专家判断高度一致,适合医疗合成数据质量评估
合成数据为解决机器学习训练中的数据稀缺问题提供了可能;然而,若缺乏质量评估,可能引入伪影、失真和不真实特征,影响模型性能与临床应用。本文提出一种基于知识的异常检测方法,用于识别合成图像中由网络生成导致的形状伪影。该方法采用两阶段框架:(i)新型特征提取器通过分析图像中解剖边界的角度梯度分布,构建专用特征空间;(ii)基于孤立森林的异常检测器。在分别基于CSAW-M和VinDr-Mammo患者数据集训练的两个合成乳腺影像数据集上验证了该方法的有效性。定量评估显示,方法成功将伪影集中在最异常的1%分区,AUC分别为0.97(CSAW-syn)和0.91(VMLO-syn)。三位影像科学家参与的阅读研究证实,算法识别出的含伪影图像,人类读者平均同意率分别为66%(CSAW-syn)和68%(VMLO-syn),约为最不异常分区的1.5–2倍。算法与人工排序的Kendall-Tau相关系数分别为0.45和0.43,表明在难以察觉的伪影检测任务中仍具合理一致性。该方法推动了合成数据的负责任使用,使开发者可依据已知解剖约束评估合成图像,并定位与修正具体问题,从而提升合成数据集的整体质量。
原文摘要 · Abstract (English)
Synthetic data provides a promising approach to address data scarcity for training machine learning models; however, adoption without proper quality assessments may introduce artifacts, distortions, and unrealistic features that compromise model performance and clinical utility. This work introduces a novel knowledge-based anomaly detection method for detecting network-induced shape artifacts in synthetic images. The introduced method utilizes a two-stage framework comprising (i) a novel feature extractor that constructs a specialized feature space by analyzing the per-image distribution of angle gradients along anatomical boundaries, and (ii) an isolation forest-based anomaly detector. We demonstrate the effectiveness of the method for identifying network-induced shape artifacts in two synthetic mammography datasets from models trained on CSAW-M and VinDr-Mammo patient datasets respectively. Quantitative evaluation shows that the method successfully concentrates artifacts in the most anomalous partition (1st percentile), with AUC values of 0.97 (CSAW-syn) and 0.91 (VMLO-syn). In addition, a reader study involving three imaging scientists confirmed that images identified by the method as containing network-induced shape artifacts were also flagged by human readers with mean agreement rates of 66% (CSAW-syn) and 68% (VMLO-syn) for the most anomalous partition, approximately 1.5-2 times higher than the least anomalous partition. Kendall-Tau correlations between algorithmic and human rankings were 0.45 and 0.43 for the two datasets, indicating reasonable agreement despite the challenging nature of subtle artifact detection. This method is a step forward in the responsible use of synthetic data, as it allows developers to evaluate synthetic images for known anatomic constraints and pinpoint and address specific issues to improve the overall quality of a synthetic dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。