arXiv:2503.15438cs.CLcs.AI2025-03被引 6

一站式蛋白工程数据与模型微调平台,打通生物与计算机科研壁垒。

VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning

  • 集成40+蛋白数据集与40+语言模型,支持多任务基准测试
  • 提供命令行与无代码界面,降低跨学科使用门槛
  • 开源可复现,适合生物、计算机双背景研究者使用

自然语言处理(NLP)已扩展至蛋白工程领域,预训练蛋白语言模型(PLMs)表现优异。然而,跨学科应用受限于数据获取难、任务评估不统一及应用复杂等问题。本文提出VenusFactory,一个融合生物数据检索、标准化任务基准测试与模块化PLM微调的统一平台。该平台支持命令行与Gradio无代码界面,集成40余个蛋白相关数据集和40余个主流PLMs。所有代码已开源,地址为https://github.com/tyang816/VenusFactory。

原文摘要 · Abstract (English)

Natural language processing (NLP) has significantly influenced scientific domains beyond human language, including protein engineering, where pre-trained protein language models (PLMs) have demonstrated remarkable success. However, interdisciplinary adoption remains limited due to challenges in data collection, task benchmarking, and application. This work presents VenusFactory, a versatile engine that integrates biological data retrieval, standardized task benchmarking, and modular fine-tuning of PLMs. VenusFactory supports both computer science and biology communities with choices of both a command-line execution and a Gradio-based no-code interface, integrating $40+$ protein-related datasets and $40+$ popular PLMs. All implementations are open-sourced on https://github.com/tyang816/VenusFactory.

蛋白工程语言模型数据平台开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。