arXiv:2509.18529cs.LGq-bio.GN2025-09被引 3

让DNA模型对序列及其反向互补序列给出一致预测。

Reverse-Complement Consistency for DNA Language Models

  • 引入反向互补一致性正则化,直接惩罚模型对原序列与反向互补序列的预测差异。
  • 在多种任务中显著减少预测翻转错误,提升鲁棒性且不牺牲准确率。
  • 适用于各类DNA模型,为生物序列分析提供可靠、高效的统一优化方案。

DNA的一个基本特性是其反向互补(RC)序列通常具有相同的生物学意义。然而,当前先进的DNA语言模型常无法捕捉这种对称性,导致对序列及其反向互补序列产生不一致的预测,影响了模型的可靠性。本文提出反向互补一致性正则化(RCCR),一种简单且模型无关的微调目标,通过直接惩罚模型在原序列与反向互补序列上的预测偏差来增强一致性。我们在三种不同骨干模型(Nucleotide Transformer、HyenaDNA、DNABERT-2)上,针对序列分类、标量回归和谱型预测等广泛基因组任务进行了评估。实验表明,RCCR显著提升了反向互补鲁棒性,大幅减少了预测翻转和错误,同时在任务准确率上优于或相当基准方法(如反向互补数据增强和测试时平均)。通过将关键生物学先验直接融入学习过程,RCCR提供了一种单一、内在鲁棒且计算高效的模型微调方案,适用于多样化的生物任务。

原文摘要 · Abstract (English)

A fundamental property of DNA is that the reverse complement (RC) of a sequence often carries identical biological meaning. However, state-of-the-art DNA language models frequently fail to capture this symmetry, producing inconsistent predictions for a sequence and its RC counterpart, which undermines their reliability. In this work, we introduce Reverse-Complement Consistency Regularization (RCCR), a simple and model-agnostic fine-tuning objective that directly penalizes the divergence between a model's prediction on a sequence and the aligned prediction on its reverse complement. We evaluate RCCR across three diverse backbones (Nucleotide Transformer, HyenaDNA, DNABERT-2) on a wide range of genomic tasks, including sequence classification, scalar regression, and profile prediction. Our experiments show that RCCR substantially improves RC robustness by dramatically reducing prediction flips and errors, all while maintaining or improving task accuracy compared to baselines such as RC data augmentation and test-time averaging. By integrating a key biological prior directly into the learning process, RCCR produces a single, intrinsically robust, and computationally efficient model fine-tuning recipe for diverse biology tasks.

DNA模型序列对称生物先验鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。