arXiv:2604.00058q-bio.GNcs.AI2026-04

GenoBERT用Transformer无参考面板精准填补基因型缺失,跨人群表现稳定。

GenoBERT: A Language Model for Accurate Genotype Imputation

  • 用Transformer直接处理分相基因型,捕捉长短程连锁不平衡
  • 在25%缺失时准确率r²达0.98,50%缺失仍保持r²>0.90
  • 无需参考面板,适合小样本和低连锁区域,适用性广

基因型填补可实现全基因组关联与风险预测研究的密集位点覆盖,但传统参考面板方法受种族偏差和稀有变异准确性限制。本文提出基于Transformer的无参考框架GenoBERT,将分相基因型进行分词,并利用自注意力机制捕捉短程与长程连锁不平衡(LD)依赖关系。在路易斯安那骨质疏松研究(LOS)和1000基因组计划(1KGP)两个独立数据集上,跨多种族群与基因型缺失水平(5%-50%)的基准测试表明,GenoBERT在整体准确率上优于四种基线方法(Beagle5.4、SCDA、BiU-Net、STICI)。在实际缺失水平(最高25%)下,其整体填补准确率可达r²≈0.98,即使在50%缺失时仍保持r²>0.90。不同族群实验验证了其一致性能提升,且对小样本和弱连锁结构具有鲁棒性。通过LD衰减分析,128个SNP(单核苷酸多态性)上下文窗口(约100 Kb)已足够捕获局部相关结构。该方法摆脱参考面板依赖,同时保持高精度,为基因型填补提供了可扩展、稳健的解决方案,也为下游基因组建模奠定基础。

原文摘要 · Abstract (English)

Genotype imputation enables dense variant coverage for genome-wide association and risk-prediction studies, yet conventional reference-panel methods remain limited by ancestry bias and reduced rare-variant accuracy. We present Genotype Bidirectional Encoder Representations from Transformers (GenoBERT), a transformer-based, reference-free framework that tokenizes phased genotypes and uses a self-attention mechanism to capture both short- and long-range linkage disequilibrium (LD) dependencies. Benchmarking on two independent datasets including the Louisiana Osteoporosis Study (LOS) and the 1000 Genomes Project (1KGP) across ancestry groups and multiple genotype missingness levels (5-50%) shows that GenoBERT achieves the highest overall accuracy compared to four baseline methods (Beagle5.4, SCDA, BiU-Net, and STICI). At practical sparsity levels (up to 25% missing), GenoBERT attains high overall imputation accuracy ($r^2 approx 0.98$) across datasets, and maintains robust performance ($r^2 > 0.90$) even at 50% missingness. Experimental results across different ancestries confirm consistent gains across datasets, with resilience to small sample sizes and weak LD. A 128-SNP (single-nucleotide polymorphism) context window (approximately 100 Kb) is validated through LD-decay analyses as sufficient to capture local correlation structures. By eliminating reference-panel dependence while preserving high accuracy, GenoBERT provides a scalable and robust solution for genotype imputation and a foundation for downstream genomic modeling.

基因组学Transformer填补算法无参考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。