arXiv:2510.18571q-bio.GNcs.LG2025-10被引 1

用多证据框架从低统计效能癌症数据中挖出真实基因信号

A Multi-Evidence Framework Rescues Low-Power Prognostic Signals and Rejects Statistical Artifacts in Cancer Genomics

  • 融合因果推断与生物验证,五重标准识别基因关联
  • 在死亡率仅13.8%的乳腺癌数据中,发现关键基因KMT2C信号弱但可信
  • 可区分假阳性(如心脏基因RYR2)与需验证的潜在驱动基因

标准癌症基因组关联研究依赖多重检验校正下的统计显著性,但在低效能队列中系统性失效。以TCGA-BRCA乳腺癌队列(n=967,133例死亡)为例,事件率仅13.8%,导致已知驱动基因出现假阴性,大段乘客基因产生假阳性。我们提出一种五准则计算框架,整合因果推断(逆概率加权、双重稳健估计)与独立生物学验证(表达、突变模式、文献证据)。标准Cox+FDR方法在该队列中未能检测到任何基因(FDR<0.05),证实其完全失效。本框架正确识别出无癌症功能的心脏基因RYR2为假阳性(名义p=0.024),同时发现尽管统计显著性边缘(p=0.047,q=0.954),KMT2C仍具强生物学证据支持。功率分析显示基因平均效能仅为15.1%,其中KMT2C仅29.8%(HR=1.55),解释了其边际显著性。通过突变模式区分真伪信号:RYR2具29.8%同义突变且无热点(乘客特征),而KMT2C仅6.7%同义突变,却有31.4%截断变异(驱动特征)。该多证据方法为低效能队列分析提供范式,强调生物可解释性优先于纯统计显著性。

原文摘要 · Abstract (English)

Motivation: Standard genome-wide association studies in cancer genomics rely on statistical significance with multiple testing correction, but systematically fail in underpowered cohorts. In TCGA breast cancer (n=967, 133 deaths), low event rates (13.8%) create severe power limitations, producing false negatives for known drivers and false positives for large passenger genes. Results: We developed a five-criteria computational framework integrating causal inference (inverse probability weighting, doubly robust estimation) with orthogonal biological validation (expression, mutation patterns, literature evidence). Applied to TCGA-BRCA mortality analysis, standard Cox+FDR detected zero genes at FDR<0.05, confirming complete failure in underpowered settings. Our framework correctly identified RYR2 -- a cardiac gene with no cancer function -- as a false positive despite nominal significance (p=0.024), while identifying KMT2C as a complex candidate requiring validation despite marginal significance (p=0.047, q=0.954). Power analysis revealed median power of 15.1% across genes, with KMT2C achieving only 29.8% power (HR=1.55), explaining borderline statistical significance despite strong biological evidence. The framework distinguished true signals from artifacts through mutation pattern analysis: RYR2 showed 29.8% silent mutations (passenger signature) with no hotspots, while KMT2C showed 6.7% silent mutations with 31.4% truncating variants (driver signature). This multi-evidence approach provides a template for analyzing underpowered cohorts, prioritizing biological interpretability over purely statistical significance. Availability: All code and analysis pipelines available at github.com/akarlaraytu/causal-inference-for-cancer-genomics

癌症基因组因果推断低效能分析多证据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。