评测大模型在生物分子多尺度建模中的极限,发现其在复杂任务上表现不足。
The limits of bio-molecular modeling with large language models : a cross-scale evaluation
- 构建跨尺度生物分子基准测试BioMol-LLM-Bench,含26项任务。
- 大模型在分类任务中表现尚可,但回归任务能力仍弱。
- 提示链数据对生物任务帮助有限,混合架构更适配长序列。
跨分子尺度的生物分子系统建模仍是科学研究的核心挑战。大语言模型(LLMs)在生物分子发现中的应用日益广泛,但针对多尺度生物问题的系统性评估及工具增强能力的严谨分析仍显不足。本文提出跨尺度生物分子基准测试BioMol-LLM-Bench,一个包含26个下游任务的统一框架,涵盖4种不同难度层级,并集成计算工具以实现更全面的评估。对13个代表性模型的评估揭示四大发现:思维链数据在生物任务中益处有限,甚至可能降低性能;混合Mamba-注意力架构更适合处理长生物分子序列;监督微调虽提升特定任务专精度,但损害通用性;当前大模型在分类任务表现良好,但在具有挑战性的回归任务中仍显薄弱。这些发现为未来基于大模型的分子系统建模提供了实践指导。
原文摘要 · Abstract (English)
The modeling of bio-molecular system across molecular scales remains a central challenge in scientific research. Large language models (LLMs) are increasingly applied to bio-molecular discovery, yet systematic evaluation across multi-scale biological problems and rigorous assessment of their tool-augmented capabilities remain limited. We reveal a systematic gap between LLM performance and mechanistic understanding through the proposed cross-scale bio-molecular benchmark: BioMol-LLM-Bench, a unified framework comprising 26 downstream tasks that covers 4 distinct difficulty levels, and computational tools are integrated for a more comprehensive evaluation. Evaluation on 13 representative models reveals 4 main findings: chain-of-thought data provides limited benefit and may even reduce performance on biological tasks; hybrid mamba-attention architectures are more effective for long bio-molecular sequences; supervised fine-tuning improves specialization at the cost of generalization; and current LLMs perform well on classification tasks but remain weak on challenging regression tasks. Together, these findings provide practical guidance for future LLM-based modeling of molecular systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。