用原型引导生成病理图像,60倍少数据也能达到顶尖模型性能。
Prototype-Guided Diffusion for Digital Pathology: Achieving Foundation Model Performance with Minimal Clinical Data
- 基于组织学原型的扩散模型生成高保真合成病理图。
- 仅需真实数据的1/60至1/760即达竞争性表现。
- 适合需要少样本训练的病理AI研发团队使用。
数字病理领域的基础模型依赖大规模数据学习复杂组织切片的紧凑特征表示。然而,数据规模与性能之间的关系缺乏透明度,引发是否必须不断增加真实数据的疑问。本文提出一种原型引导的扩散模型,可大规模生成高质量合成病理数据,支持大范围自监督学习,减少对真实患者样本的依赖,同时保持下游任务性能。在采样过程中引入组织学原型指导,确保生成数据具有生物学和诊断意义的变异。实验表明,基于合成数据训练的自监督特征模型,在仅使用真实数据1/60至1/760的情况下,仍能实现竞争性表现;在多个评估指标和任务中,其性能与使用数倍甚至数十倍更大真实数据集训练的模型相当或更优。结合合成与真实数据的混合策略进一步提升了性能,在多项评估中取得最佳结果。这些发现证明生成式AI可高效构建数字病理训练数据,显著降低对临床数据的依赖,凸显本方法的效率优势。
原文摘要 · Abstract (English)
Foundation models in digital pathology use massive datasets to learn useful compact feature representations of complex histology images. However, there is limited transparency into what drives the correlation between dataset size and performance, raising the question of whether simply adding more data to increase performance is always necessary. In this study, we propose a prototype-guided diffusion model to generate high-fidelity synthetic pathology data at scale, enabling large-scale self-supervised learning and reducing reliance on real patient samples while preserving downstream performance. Using guidance from histological prototypes during sampling, our approach ensures biologically and diagnostically meaningful variations in the generated data. We demonstrate that self-supervised features trained on our synthetic dataset achieve competitive performance despite using ~60x-760x less data than models trained on large real-world datasets. Notably, models trained using our synthetic data showed statistically comparable or better performance across multiple evaluation metrics and tasks, even when compared to models trained on orders of magnitude larger datasets. Our hybrid approach, combining synthetic and real data, further enhanced performance, achieving top results in several evaluations. These findings underscore the potential of generative AI to create compelling training data for digital pathology, significantly reducing the reliance on extensive clinical datasets and highlighting the efficiency of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。