数据打乱方式影响基因语言模型评测结果,4%差异来自硬件设置
Same model, better performance: the impact of shuffling on DNA Language Models benchmarking
- 预打乱数据再存储,消除硬件依赖带来的评测偏差
- 相同模型在不同数据加载设置下性能差达4%,影响排名
- 适合基因组学与机器学习交叉研究者参考
大语言模型在基因组学中应用日益广泛,亟需标准化基准评估其能力。然而,评估DNA语言模型(DNA LMs)涉及基因组领域特有挑战与机器学习方法的交叉,细微实现细节可能严重损害基准有效性。我们通过BEND(Benchmarking DNA Language Models)揭示:硬件相关的超参数——数据加载工作数和缓冲区大小——会导致相同模型性能出现高达4%的虚假波动。问题根源在于数据打乱不足与基因组数据特性相互作用。对HyenaDNA、DNABERT-2、ResNet-LM三个模型的实验表明,此类伪影同时影响绝对性能和相对模型排名。我们提出简单解决方案:存储前预打乱数据,可消除硬件依赖且保持效率。本研究揭示标准机器学习实践可能与领域特异性数据产生意外交互,对特殊领域基准设计具有广泛启示。
原文摘要 · Abstract (English)
Large Language Models are increasingly popular in genomics due to their potential to decode complex biological sequences. Hence, researchers require a standardized benchmark to evaluate DNA Language Models (DNA LMs) capabilities. However, evaluating DNA LMs is a complex task that intersects genomic's domain-specific challenges and machine learning methodologies, where seemingly minor implementation details can significantly compromise benchmark validity. We demonstrate this through BEND (Benchmarking DNA Language Models), where hardware-dependent hyperparameters -- number of data loading workers and buffer sizes -- create spurious performance variations of up to 4% for identical models. The problem stems from inadequate data shuffling interacting with domain specific data characteristics. Experiments with three DNA language models (HyenaDNA, DNABERT-2, ResNet-LM) show these artifacts affect both absolute performance and relative model rankings. We propose a simple solution: pre-shuffling data before storage eliminates hardware dependencies while maintaining efficiency. This work highlights how standard ML practices can interact unexpectedly with domain-specific data characteristics, with broader implications for benchmark design in specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。