构建前列腺癌变异分类基准,提升不确定意义变异的临床判读准确率。
Prostate-VarBench: A Benchmark with Interpretable TabNet Framework for Prostate Cancer Variant Classification
- 整合三大数据库构建19.3万例前列腺特异性变异数据集,支持患者/基因级划分防泄露。
- 模型达89.9%准确率,修正注释缺陷后不确定变异占比下降6.5个百分点。
- 采用可解释的TabNet模型,输出结果符合临床肿瘤委员会判读逻辑,适合医疗科研人员使用。
不确定意义的变异(VUS)因缺乏致病性或良性证据,限制了前列腺癌基因组学的临床应用,导致诊断与治疗延迟。现有研究受限于不同来源标注不一致,且缺乏前列腺特异性基准用于公平比较。本文提出Prostate-VarBench,一个定制化管道,整合COSMIC(体细胞突变)、ClinVar(专家标注临床变异)和TCGA-PRAD(癌症基因组图谱前列腺肿瘤基因组数据),构建包含193,278个变异的标准化数据集,并支持患者或基因级划分以防止数据泄露。为保障数据质量,修复了影响临床意义字段的VEP工具问题,该问题曾将多个转录本记录错误合并。进一步统一了56个可解释特征,覆盖八个临床相关层级,包括种群频率、变异类型及临床背景。引入AlphaMissense致病性评分以增强错义变异分类,降低VUS不确定性。基于此资源,训练了一个可解释的TabNet模型,其分步稀疏掩码能生成与分子肿瘤委员会评审实践一致的个体化推理依据。在独立测试集上,模型达到89.9%准确率,且平衡类指标表现良好;经VEP修正后,VUS比例绝对下降6.5%。
原文摘要 · Abstract (English)
Variants of Uncertain Significance (VUS) limit the clinical utility of prostate cancer genomics by delaying diagnosis and therapy when evidence for pathogenicity or benignity is incomplete. Progress is further limited by inconsistent annotations across sources and the absence of a prostate-specific benchmark for fair comparison. We introduce Prostate-VarBench, a curated pipeline for creating prostate-specific benchmarks that integrates COSMIC (somatic cancer mutations), ClinVar (expert-curated clinical variants), and TCGA-PRAD (prostate tumor genomics from The Cancer Genome Atlas) into a harmonized dataset of 193,278 variants supporting patient- or gene-aware splits to prevent data leakage. To ensure data integrity, we corrected a Variant Effect Predictor (VEP) issue that merged multiple transcript records, introducing ambiguity in clinical significance fields. We then standardized 56 interpretable features across eight clinically relevant tiers, including population frequency, variant type, and clinical context. AlphaMissense pathogenicity scores were incorporated to enhance missense variant classification and reduce VUS uncertainty. Building on this resource, we trained an interpretable TabNet model to classify variant pathogenicity, whose step-wise sparse masks provide per-case rationales consistent with molecular tumor board review practices. On the held-out test set, the model achieved 89.9% accuracy with balanced class metrics, and the VEP correction yields an 6.5% absolute reduction in VUS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。