arXiv:2509.12266q-bio.GNcs.LG2025-09

一站式工具库,简化基因组模型训练、部署与解释流程。

Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models

  • 整合数据收集、模型微调、推理与可解释性分析全流程
  • 支持多种模型与微调方法,在两个公开基准上验证性能
  • 适合基因组研究者快速构建和分析基础模型

我们提出Genome-Factory,首个用于基因组基础模型调优、部署与解释的集成Python库。其核心在于统一基因组模型开发工作流:数据收集、模型微调、推理、基准测试与可解释性分析。数据收集方面,提供自动化管道下载并预处理基因组序列;模型微调支持全量与参数高效微调,适配多种基因组模型;推理支持嵌入提取与DNA序列生成;基准测试包含两个现有基准,并提供灵活接口扩展;可解释性方面,引入基于稀疏自编码器的开源生物学解释器。我们在三个维度验证其有效性:(i) 兼容多种模型与微调方法;(ii) 使用两个开源基准评估下游性能;(iii) 对DNABERT-2学习表征进行生物解释。结果表明其在真实基因组分析中的实用价值。GitHub: https://github.com/WeiminWu2000/Genome_Factory。

原文摘要 · Abstract (English)

We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. Our core contribution is to simplify and unify the workflow for genomic model development: data collection, model tuning, inference, benchmarking, and interpretability. For data collection, Genome-Factory offers an automated pipeline to download genomic sequences and preprocess them. For model tuning, Genome-Factory supports both full and parameter-efficient fine-tuning across diverse genomic models. For inference, Genome-Factory enables both embedding extraction and DNA sequence generation. For benchmarking, we include two existing benchmarks and provide a flexible interface to incorporate additional benchmarks. For interpretability, Genome-Factory introduces an open-source biological interpreter based on a sparse auto-encoder. We validate the utility of Genome-Factory across three dimensions: (i) Compatibility with diverse models and fine-tuning methods; (ii) Benchmarking downstream performance using two open-source benchmarks; (iii) Biological interpretation of learned representations with DNABERT-2. These results highlight its practical value for real-world genomic analysis. GitHub: https://github.com/WeiminWu2000/Genome_Factory.

基因组模型模型部署可解释性Python库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。