arXiv:2506.22963stat.MLcs.LG2025-06

CN-SBM通过分块建模揭示癌症拷贝数变异的主成分与残差模式。

CN-SBM: Categorical Block Modelling For Primary and Residual Copy Number Variation

  • 用双阶段分块模型分离拷贝数变异的主成分与细粒度异常
  • 在真实数据中比现有方法更优,发现临床相关亚型
  • 适合研究肿瘤异质性与预后分层的科研人员

癌症是基因组紊乱疾病,其克隆演化可通过追踪全基因组拷贝数变异(CNV)来监测。本文提出拷贝数随机块模型(CN-SBM),一种基于二部分类块模型的统计框架,可同时对样本和基因组区域进行聚类,依据离散拷贝数状态划分。与依赖高斯或泊松假设的模型不同,CN-SBM尊重拷贝数变异的离散特性,并通过块结构捕捉亚群特异性模式。采用两阶段方法,将CNV数据分解为主成分与残差成分,从而检测大规模染色体改变及细微异常。我们推导出适用于大规模队列与高分辨率数据的可扩展变分推断算法。在模拟与真实数据集上的基准测试显示,该模型拟合效果优于现有方法。应用于TCGA低级别胶质瘤数据,CN-SBM揭示了具有临床意义的亚型及结构化残差变异,有助于生存分析中的患者分层。这些结果确立了CN-SBM作为可解释、可扩展的拷贝数变异分析框架,对肿瘤异质性与预后研究具有直接价值。

原文摘要 · Abstract (English)

Cancer is a genetic disorder whose clonal evolution can be monitored by tracking noisy genome-wide copy number variants. We introduce the Copy Number Stochastic Block Model (CN-SBM), a probabilistic framework that jointly clusters samples and genomic regions based on discrete copy number states using a bipartite categorical block model. Unlike models relying on Gaussian or Poisson assumptions, CN-SBM respects the discrete nature of CNV calls and captures subpopulation-specific patterns through block-wise structure. Using a two-stage approach, CN-SBM decomposes CNV data into primary and residual components, enabling detection of both large-scale chromosomal alterations and finer aberrations. We derive a scalable variational inference algorithm for application to large cohorts and high-resolution data. Benchmarks on simulated and real datasets show improved model fit over existing methods. Applied to TCGA low-grade glioma data, CN-SBM reveals clinically relevant subtypes and structured residual variation, aiding patient stratification in survival analysis. These results establish CN-SBM as an interpretable, scalable framework for CNV analysis with direct relevance for tumor heterogeneity and prognosis.

拷贝数变异生物统计肿瘤异质性块模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。