arXiv:2609.04861cs.LG2026-09

GenDA模型能精准预测基因变异影响,但生成功能序列能力弱,提示不同任务需独立验证。

When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation

论文配图:When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation
图 1 · 摘自论文原文
  • 用熵引导的跨度放置提升序列重建压力,聚焦复杂区域
  • 临床变异预测AUROC达0.774,优于自回归模型0.103
  • 零样本功能补全测试失败,生成效果不如随机打乱控制组

双向离散扩散模型天然适合基因组建模,因其可从两侧重建缺失序列。我们提出GenDA(基因组密度优化吸收扩散)模型,假设熵引导的跨度放置能集中重建压力于序列复杂区域,从而同时提升变异效应预测与功能序列生成。结果部分支持该假设:经微调后,202M参数的GenDA模型在综合ClinVar SNV数据上获得0.774的AUROC,比同等规模自回归模型高0.103。然而,匹配的随机跨度变异也达到0.777,未显示熵引导带来优势。更意外的是,GenDA在零样本功能补全测试中表现不佳:在启动子、增强子、外显子边界和内含子边界上,其性能未显著优于仅打乱原缺口但严格保留3-聚体组成的对照组。50–500 bp缺口时已出现失败,增强子退化随缺口变长加剧。诊断发现:熵反映局部序列复杂性而非功能重要性;1-元标记限制物理上下文;训练跨度上限为300 bp;高绝对AlphaGenome保真度可与负控制归一化恢复共存。结果表明,强微调变异预测、合理噪声先验与功能生成是独立命题,需分别验证。

原文摘要 · Abstract (English)

Bidirectional discrete diffusion model appears naturally suited to genomic modeling because it can reconstruct missing sequence from both flanks. We developed GenDA (Genomic Density-optimized Absorbing Diffusion) under the additional hypothesis that entropy-guided span placement would concentrate reconstruction pressure on compositionally complex regions, improving both downstream variant-effect prediction and functional sequence generation. Our results only partially support this premise. After supervised fine-tuning, the 202M-parameter GenDA model reaches a pooled ClinVar SNV AUROC of 0.774, exceeding a similarly scaled autoregressive model by 0.103. However, a matched random-span variant reaches 0.777, providing no evidence that entropy guidance causes the ClinVar improvement. More unexpectedly, GenDA fails a zero-shot functional inpainting stress test: across promoters, enhancers, exon boundaries, and intron boundaries, it does not consistently outperform a control that shuffles the native gap while exactly preserving 3-mer composition. Failure is already present for 50--500-bp gaps, although enhancer degradation worsens at longer gaps. Diagnostics identify several boundary conditions: entropy measures local sequence complexity rather than functional importance; 1-mer tokenization limits physical context; training spans are capped at 300 bp; and high absolute AlphaGenome fidelity can coexist with negative control-normalized restoration. These results show that strong fine-tuned variant prediction, a plausible corruption prior, and functional generation are distinct claims that require separate validation.

基因组建模扩散模型功能生成变异预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。