arXiv:2607.27712cs.LG2026-07

根据古DNA损伤位置分布,动态调整掩码策略,显著提升重建精度。

VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers, Validated on Ancient DNA Reconstruction

  • 基于实测损伤图谱,按位置分布智能分配掩码概率。
  • 在古DNA重建中,准确率提升4.18至10.35个百分点,交叉熵降低13%。
  • 适用于各类序列降解场景,可推广至病理、单细胞等噪声数据训练。

标准掩码语言模型微调对所有位置采用统一掩码率,假设重构难度与位置无关。但当降解过程具有可预测的位置特异性时,该假设失效:在损伤高峰区域,模型性能甚至低于频率匹配的随机预测器。本文提出VESTIGE,一种无需参数、可直接替换标准MLM的数据处理方法,其掩码分布与实测的每位置损伤率对齐。应用于古DNA(aDNA)重建任务,利用mapDamage2量化胞嘧啶脱氨导致的C→T/G→A梯度。将平均C/G掩码率重设为15%(与标准MLM一致),以空间重分布为唯一变量,在猛犸象编码区数据集(两样本,七基因)上,使用DNABERT-2模型进行对比实验。在六个终端区宽度和626个窗口中,VESTIGE在所有宽度下均优于标准MLM(提升4.18至10.35个百分点,所有p<10^-8),验证集交叉熵下降13%(3.274 vs. 3.757),ESMFold重构的TM分数全部高于0.95,即使在损伤率放大10-30倍的情况下仍保持高精度。1D CNN生物安全分类器实现AUC=0.935,清除98.2%的重建窗口,剩余1.76%归因于参考基因组特征,而非重构伪影。该原则具有领域普适性:任何可测量的位置或上下文相关损伤模式(如FFPE、亚硫酸氢盐、宏基因组、纳米孔测序)均可替代PMD数组,使VESTIGE成为面向降解或噪声序列输入的智能系统知识引导训练范式。

原文摘要 · Abstract (English)

Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is position-agnostic. When the degradation process is characterised and concentrated at predictable positions, this assumption fails: at peak damage sites the model can underperform a frequency-matched random predictor. We introduce VESTIGE, a parameter-free, drop-in replacement for the standard MLM collator that aligns the masking distribution with an empirically measured per-position corruption profile. We apply it to ancient DNA (aDNA) reconstruction, where cytosine deamination produces a position-dependent C-to-T / G-to-A gradient quantified per-position by mapDamage2. Rescaling so the mean C/G masking rate equals 15% - identical to standard MLM - isolates spatial redistribution as the sole variable, with model, data, seed, and hyperparameters held fixed across both DNABERT-2 runs on a mammoth CDS corpus (two specimens, seven genes). Across six terminal-zone widths and 626 paired windows, VESTIGE leads standard MLM at every width (Delta = +4.18 to +10.35 pp, all p < 10^-8), cuts validation cross-entropy by 13% (3.274 vs. 3.757), and yields ESMFold reconstructions with TM-score > 0.95 across all six reconstructions (three genes) even under damage amplified 10-30x beyond authentic PMD rates. A 1D CNN biosecurity classifier returns AUC = 0.935 and clears 98.2% of reconstructed windows, the 1.76% remainder attributable to reference-genome features, not reconstruction artefacts. The principle is domain-agnostic: any measurable position- or context-specific corruption profile - FFPE, bisulfite, metagenomic, or nanopore - substitutes directly for the PMD array, making VESTIGE a knowledge-guided training routine for intelligent systems operating on degraded or noisy sequence inputs.

古基因组掩码策略序列重建降解数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。