综述2022-2025年无创多癌早筛中cfDNA分析的计算方法与挑战
Computational Methods and Challenges in Cell-Free DNA Analysis for Multi-Cancer Early Detection

- 结合片段组学与表观遗传特征,用机器学习和深度学习识别早期癌症信号
- 多模态集成方法在临床转化中最具潜力,但需统一评估标准
- 适合关注液体活检、生物信息算法及癌症早筛研究者阅读
细胞游离DNA(cfDNA)为无创多癌早筛(MCED)提供了有前景的路径,可通过一次血液检测同时筛查多种癌症,尤其对缺乏现有筛查方案的癌症具有高灵敏度。本文综述2022至2025年间针对cfDNA的计算方法,聚焦片段组学与表观遗传特征的提取与分析策略以实现早期癌症检测。首先简要介绍cfDNA信号的生物学基础,随后回顾经典统计方法、机器学习模型以及基于自编码器的深度学习框架。针对每种方法,讨论其生物可解释性、验证策略及临床整合准备度。此外,将当前挑战分为技术、计算与方法三类,并指出领域内未解问题。结果显示,多模态集成方法在临床应用中最具前景且准备度最高;然而,未来研究的评估与横向比较亟需标准化评价协议与结果报告。
原文摘要 · Abstract (English)
Cell-free DNA (cfDNA) is a promising avenue for non-invasive multicancer early detection (MCED), in that, it can enable multiple cancer detection simultaneously from a single blood draw, with particular sensitivity to cancers that currently lack established screening programs. Here we review the computational methods developed between 2022 and 2025 for cfDNA-based MCED. We focus on how fragmentomics and epigenetic features are extracted and analyzed to detect cancer at early stages. We first briefly outline the biological basis of cfDNA signals, then review classical statistical and machine learning approaches alongside deep learning frameworks including autoencoder-based models. For each method we discuss biological interpretability, validation strategy, and readiness for clinical integration. Furthermore, we categorize the current challenges into technical, computational, and methodological while outlining open problems in the field. This review shows that multimodal ensemble approaches have the strongest promise for clinical integration and the highest readiness. However, for better assessment of future work and side-by-side comparison, standardization of evaluation protocols and reporting results will be crucial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。