arXiv:2411.01025cs.CV2024-11

用合成数据训练模型,提升癌症基因异常检测的准确率与不确定性判断能力

FISHing in Uncertainty: Synthetic Contrastive Learning for Genetic Aberration Detection

  • 通过合成图像替代人工标注,降低数据成本
  • 在真实数据上实现96.7%准确率(最确定50%样本)
  • 适合临床诊断中需要高可靠性与自动化的需求

基因异常检测对癌症诊断至关重要,通常依赖荧光原位杂交(FISH)。现有方法因信号变异、依赖昂贵的人工标注且未能有效处理内在不确定性而受限。本文提出一种新方法:利用合成图像消除人工标注需求,并采用联合对比学习与分类目标进行训练,以有效应对类别间差异。我们在真实世界FISH图像的人工标注数据集上验证了该方法的优越泛化能力与不确定性校准性能。模型在最确定的50%样本中达到96.7%的分类准确率,兼具高准确性与可靠不确定性估计。所提端到端方法显著降低人力与时间成本,提升诊断效率。所有代码与数据公开于:https://github.com/SimonBon/FISHing

原文摘要 · Abstract (English)

Detecting genetic aberrations is crucial in cancer diagnosis, typically through fluorescence in situ hybridization (FISH). However, existing FISH image classification methods face challenges due to signal variability, the need for costly manual annotations and fail to adequately address the intrinsic uncertainty. We introduce a novel approach that leverages synthetic images to eliminate the requirement for manual annotations and utilizes a joint contrastive and classification objective for training to account for inter-class variation effectively. We demonstrate the superior generalization capabilities and uncertainty calibration of our method, which is trained on synthetic data, by testing it on a manually annotated dataset of real-world FISH images. Our model offers superior calibration in terms of classification accuracy and uncertainty quantification with a classification accuracy of 96.7% among the 50% most certain cases. The presented end-to-end method reduces the demands on personnel and time and improves the diagnostic workflow due to its accuracy and adaptability. All code and data is publicly accessible at: https://github.com/SimonBon/FISHing

基因检测合成数据不确定性量化医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。