arXiv:2606.16540q-bio.QMcs.LG2026-06

构建可复现的生物序列模型生态,统一接口与数据标准

MultiMolecule: a modular ecosystem for biomolecular sequence-model workflows

  • 将不同生物序列模型整合为可执行的标准化流程
  • 提供53个完整模型家族和112个标准化检查点
  • 适合需要模型复现与跨任务对比的研究者

生物分子序列模型在原始研究之外被广泛重用,但公开的检查点通常缺乏执行上下文,难以追溯行为来源、适配新检测任务、在统一任务定义下比较模型或部署生物预测。MultiMolecule 是一个开源 Python 生态系统,将异构的 RNA、DNA 和蛋白质序列模型发布转化为完整的、可溯源的模型家族实现,采用统一的加载、工作流和预测接口。本资源包含 53 个完整的模型家族实现,共 112 个标准化模型检查点,以及通过 39 个公共数据集仓库发布的 16 个精选数据集资源和 10 个面向用户的预测流水线。标准化组件均关联源代码溯源、转换/准备代码、源参考校验、扩展数据摘要和公开文档,使用户可审查标准化内容、验证行为一致性,并了解各组件在训练、评估、推理或部署中的角色。通过将重用从特定仓库的检查点转向连接标准化检查点、精选数据集、运行器工作流和生物预测流水线的可执行实现,MultiMolecule 提供了保持源定义模型行为、适应新检测任务、实现可控评估和部署生物预测的通用基础设施。

原文摘要 · Abstract (English)

Biomolecular sequence models are increasingly reused outside the studies in which they were introduced, but public checkpoints rarely preserve the execution context needed to inspect source-defined behavior, adapt models to new assays, compare models under shared task definitions or deploy biological predictions. MultiMolecule is an open-source Python ecosystem that turns heterogeneous RNA, DNA and protein sequence-model releases into complete, source-checked model-family implementations with shared loading, workflow and prediction interfaces. The Resource state reported here includes 53 complete model-family implementations with 112 standardized model checkpoints, together with 16 curated dataset resources released through 39 public dataset repositories and 10 user-facing prediction pipelines. Standardized components are linked to source provenance, conversion or preparation code, source-reference checks, Extended Data summaries and public documentation, allowing users to inspect what was standardized, what behavior was checked and how each component enters training, evaluation, inference or deployment. By shifting reuse from repository-specific checkpoints to executable implementations connected to standardized checkpoints, curated datasets, Runner workflows and biological prediction pipelines, MultiMolecule provides common infrastructure for preserving source-defined model behavior, adapting models to new assays, enabling controlled evaluation and deploying biomolecular predictions.

生物序列模型复现标准化Python生态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。