首个带功能注释的抗体指令数据集,让大模型读懂并设计抗体。
AFD-INSTRUCTION: A Comprehensive Antibody Instruction Dataset with Functional Annotations for LLM-Based Understanding and Design
- 构建抗体序列与功能描述的对齐数据,支持自然语言理解与生成。
- 在多种任务上显著提升通用大模型的抗体相关性能。
- 适合抗体设计、药物发现及LLM应用研究者使用。
大语言模型(LLMs)在蛋白质表示学习方面取得了显著进展,但其通过自然语言理解与设计抗体的能力仍受限。为解决此问题,我们提出AFD-Instruction,首个专为抗体设计的大规模指令数据集,包含两大核心组件:抗体理解(从序列推断功能属性)与抗体设计(在功能约束下生成新序列)。该数据集提供明确的序列-功能对齐,支持基于自然语言指令的抗体设计。在通用大模型上的广泛指令微调实验表明,AFD-Instruction在多种抗体相关任务中持续提升性能。通过将抗体序列与功能文本描述关联,该数据集为推进抗体建模和加速治疗性抗体发现奠定了新基础。
原文摘要 · Abstract (English)
Large language models (LLMs) have significantly advanced protein representation learning. However, their capacity to interpret and design antibodies through natural language remains limited. To address this challenge, we present AFD-Instruction, the first large-scale instruction dataset with functional annotations tailored to antibodies. This dataset encompasses two key components: antibody understanding, which infers functional attributes directly from sequences, and antibody design, which enables de novo sequence generation under functional constraints. These components provide explicit sequence-function alignment and support antibody design guided by natural language instructions. Extensive instruction-tuning experiments on general-purpose LLMs demonstrate that AFD-Instruction consistently improves performance across diverse antibody-related tasks. By linking antibody sequences with textual descriptions of function, AFD-Instruction establishes a new foundation for advancing antibody modeling and accelerating therapeutic discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。