arXiv:2605.24520q-bio.GNcs.LG2026-05

构建全基因组错义变异致病性预测框架,整合多源数据提升临床解读准确率。

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

论文配图:AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction
图 1 · 摘自论文原文
  • 融合303项特征,包括进化保守、蛋白语言模型等,系统标注错义变异
  • 在13万+临床标注数据上实现0.941的MCC与0.995的AUC,表现最优
  • 提供可复现代码与全基因组预测结果,适合遗传病研究与临床辅助诊断

错义变异的致病性判断仍具挑战,因其依赖来自人群频率、进化保守性、转录本背景、氨基酸替换严重程度、已有预测工具及蛋白语言模型衍生特征的多重证据。本文提出AnnotateMissense,一个可扩展的错义变异注释、基准测试与全基因组预测框架。该框架整合hg38参考基因组上的错义变异(源自dbNSFP v5.1),结合ANNOVAR注释、dbNSFP转录本/蛋白质描述符、AlphaMissense评分、ESM衍生特征、保守性度量、人群频率变量、经典致病性预测器及人工设计的氨基酸/密码子上下文特征。基于132,714个ClinVar标注的错义变异,我们在受控特征配置下对机器学习与深度学习模型进行基准测试。包含全部303个特征的完整集使用XGBoost达到最佳性能:分层五折交叉验证中平均MCC为0.9411,ROC-AUC为0.9950。受限的朴素与位置导向特征集最佳MCC分别为0.4989和0.5113。环形控制消融实验表明,移除先验预测器、人群频率及临床重叠证据会降低性能,而单独排除AlphaMissense与ESM衍生特征影响较小。时间序列ClinVar验证在新发现的致病/良性变异上取得MCC=0.7613,准确率=0.8798,F1分数=0.8750。最终模型应用于90,643,830个hg38错义变异,生成致病性评分与二分类标签。代码与输出已公开于https://github.com/MuhammadMuneeb007/CAGI7_Annotate_All_Missense 和 https://doi.org/10.5281/zenodo.19981867。

原文摘要 · Abstract (English)

Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary conservation, transcript context, amino acid substitution severity, prior pathogenicity predictors and protein-language-model-derived features. We present AnnotateMissense, a scalable annotation, benchmarking and genome-wide prediction framework for missense variant interpretation. AnnotateMissense integrates hg38 missense variants derived from dbNSFP v5.1 with ANNOVAR annotations, dbNSFP transcript/protein descriptors, AlphaMissense scores, ESM-derived features, conservation metrics, population-frequency variables, established pathogenicity predictors and engineered amino acid/codon-context features. Using 132,714 ClinVar-labelled missense variants, we benchmarked machine-learning and deep-learning models under controlled feature configurations. The full 303-feature benchmark set achieved the strongest performance with XGBoost, reaching mean MCC = 0.9411 and ROC-AUC = 0.9950 across stratified five-fold cross-validation. Restricted naive and location-oriented feature sets achieved lower best MCC values of 0.4989 and 0.5113, respectively. Circularity-controlled ablations showed that removing prior-predictor, population-frequency and clinically overlapping evidence reduced performance, whereas excluding AlphaMissense and ESM-derived features alone had minimal effect. Temporal ClinVar validation on newly observed pathogenic/benign variants achieved MCC = 0.7613, accuracy = 0.8798 and F1-score = 0.8750. The final model was applied to 90,643,830 hg38 missense variants to generate AnnotateMissense pathogenicity scores and binary prediction labels. Code and outputs are available at https://github.com/MuhammadMuneeb007/CAGI7_Annotate_All_Missense and https://doi.org/10.5281/zenodo.19981867.

基因组注释致病性预测机器学习错义变异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。