CHROMA模型可自动识别染色体异常,助力精准肿瘤诊疗。
An Inclusive Foundation Model for Generalizable Cytogenetics in Precision Oncology
- 基于自监督学习,用超8.4万样本预训练,学会通用染色体异常表征。
- 在少标注、不平衡数据下仍优于其他方法,检测各类异常更准确。
- 适合临床科研与病理专家,推动罕见基因异常早筛和AI普适应用。
染色体分析对诊断遗传病和指导癌症治疗至关重要,依赖于对体细胞克隆异常的识别。然而,由于染色体异常的复杂性和多样性,构建AI模型面临巨大挑战,需大量标注工作;现有自动化方法多为任务特定,缺乏泛化能力,受限于覆盖不同资源条件的全面数据集。本文提出CHROMA,一种用于细胞基因组学的基底模型,通过在超过84,000份样本(约400万张染色体图像)上进行自监督预训练,学习染色体异常的通用表示。在各类异常检测中,即使在标注数据较少、数据分布不均衡的情况下,CHROMA仍优于其他方法。该模型可全面映射不同异常类型的基因组不稳定性及克隆病变,提供可扩展、泛化的可靠自动化临床分析方案,显著降低专家标注负担,推动罕见基因异常早期发现,促进精准肿瘤学发展,并拓展广泛临床AI应用,使高级基因组分析更易获取。
原文摘要 · Abstract (English)
Chromosome analysis is vital for diagnosing genetic disorders and guiding cancer therapy decisions through the identification of somatic clonal aberrations. However, developing an AI model are hindered by the overwhelming complexity and diversity of chromosomal abnormalities, requiring extensive annotation efforts, while automated methods remain task-specific and lack generalizability due to the scarcity of comprehensive datasets spanning diverse resource conditions. Here, we introduce CHROMA, a foundation model for cytogenomics, designed to overcome these challenges by learning generalizable representations of chromosomal abnormalities. Pre-trained on over 84,000 specimens (~4 million chromosomal images) via self-supervised learning, CHROMA outperforms other methods across all types of abnormalities, even when trained on fewer labelled data and more imbalanced datasets. By facilitating comprehensive mapping of instability and clonal leisons across various aberration types, CHROMA offers a scalable and generalizable solution for reliable and automated clinical analysis, reducing the annotation workload for experts and advancing precision oncology through the early detection of rare genomic abnormalities, enabling broad clinical AI applications and making advanced genomic analysis more accessible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。