构建首个针对调控性DNA的综合评估基准,检验大模型真实能力
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA
- 设计三类任务测试模型在零样本、探针和微调下的表现
- 多数现有模型性能平庸,且耗能远高于传统基线方法
- 适合基因组学研究者评估模型对调控元件的理解能力
近期自监督学习在自然语言、视觉和蛋白质序列领域的进展,推动了大规模基因组DNA语言模型(DNALMs)的发展。这些模型旨在学习多样化DNA元件的通用表征,有望支持基因组预测、解释与设计等任务。然而,现有评估基准未能充分衡量DNALMs在关键下游应用中的表现,尤其针对对基因调控至关重要的非编码DNA元件。本文提出DART-Eval,一套专注于调控性DNA的代表性评估基准,涵盖零样本、探针和微调场景,并以当前主流的从头建模方法为基线进行对比。评估任务包括功能序列特征发现、细胞类型特异性调控活性预测以及遗传变异影响的反事实推断。结果显示,当前DNALMs表现不一致,多数任务上未显著优于基线模型,但计算开销显著更高。我们进一步探讨了下一代DNALMs在建模、数据整理和评估方面的潜在改进方向。代码已开源:https://github.com/kundajelab/DART-Eval。
原文摘要 · Abstract (English)
Recent advances in self-supervised models for natural language, vision, and protein sequences have inspired the development of large genomic DNA language models (DNALMs). These models aim to learn generalizable representations of diverse DNA elements, potentially enabling various genomic prediction, interpretation and design tasks. Despite their potential, existing benchmarks do not adequately assess the capabilities of DNALMs on key downstream applications involving an important class of non-coding DNA elements critical for regulating gene activity. In this study, we introduce DART-Eval, a suite of representative benchmarks specifically focused on regulatory DNA to evaluate model performance across zero-shot, probed, and fine-tuned scenarios against contemporary ab initio models as baselines. Our benchmarks target biologically meaningful downstream tasks such as functional sequence feature discovery, predicting cell-type specific regulatory activity, and counterfactual prediction of the impacts of genetic variants. We find that current DNALMs exhibit inconsistent performance and do not offer compelling gains over alternative baseline models for most tasks, while requiring significantly more computational resources. We discuss potentially promising modeling, data curation, and evaluation strategies for the next generation of DNALMs. Our code is available at https://github.com/kundajelab/DART-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。