arXiv:2606.04525cs.CLcs.LG2026-06中稿 · ICML

构建基因组模型统一评估基准,揭示现有比较方式的缺陷。

GENEB: Why Genomic Models Are Hard to Compare

论文配图:GENEB: Why Genomic Models Are Hard to Compare
图 1 · 摘自论文原文
  • 用统一探针法评测40个基因组模型在100项任务上的表现
  • 发现模型排名随任务类别剧烈变化,规模提升效果有限且不一致
  • 适合关注基因组模型真实性能与选型的科研人员

基因组基础模型的发展难以评估,原因在于基准碎片化、评价协议不兼容及任务特异性报告。为此,我们提出GENEB,一个大规模诊断性基准,基于统一探针协议,在13种功能类别下的100项任务中,评估40个基因组基础模型的冻结表征,涵盖少样本场景。GENEB支持对模型规模、架构、分词方式和预训练数据的可控对比,并明确揭示任务层面的权衡。分析显示,综合排行榜极不稳定:模型排名在不同任务类别间波动剧烈;规模提升仅带来微弱且不一致的收益;架构与预训练对齐常优于参数量。这些结果暴露了当前评估方法的局限性,确立GENEB作为基因组机器学习中严谨比较与类别感知选型的参考框架。

原文摘要 · Abstract (English)

Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting. As a result, claims of superiority or generality across models are often not directly comparable. We introduce GENEB, a large-scale diagnostic benchmark that evaluates frozen representations from 40 genomic foundation models across 100 tasks spanning 13 functional categories under a unified probing-based protocol, including few-shot regimes. GENEB enables controlled comparison across model scale, architecture, tokenization, and pretraining data while explicitly exposing task-level trade-offs. Our analysis shows that aggregate leaderboards are unstable: model rankings vary sharply across task categories, scale provides only modest and inconsistent gains, and architectural and pretraining alignment frequently outweigh parameter count. These results highlight limitations of current evaluation practices and position GENEB as a reference framework for principled comparison and category-aware model selection in genomic machine learning.

基因组模型评估基准模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。