arXiv:2502.06634cs.LGcs.AI2025-02NAACL被引 1

用大模型自动扩充分子与文本的标注数据,显著提升药物发现效率

Automatic Annotation Augmentation Boosts Translation between Molecules and Natural Language

  • 基于大模型自动生成多样化的分子描述,增强数据多样性
  • 在分子生成任务中性能提升最高达301%
  • 适用于文本、图像、图结构等多种下游任务

近年来,人工智能在生物研究中的进展聚焦于整合分子数据与自然语言以加速药物发现。然而,高质量标注数据稀缺限制了该领域的发展。本文提出LA$^3$——一种基于语言的自动标注增强框架,利用大语言模型扩充现有数据集,从而提升AI训练效果。我们通过系统重写已有的分子注释,构建了增强数据集LaChEBI-20,新注释在保留关键分子信息的同时,引入更丰富的句式与词汇。基于基准架构训练的LaMolT5模型在基于文本的从头分子生成和分子图文描述任务中表现优异,超越现有最优模型。实验表明,引入LA$^3$后性能最高提升301%。此外,我们在图像、文本和图任务中验证了LA$^3$的有效性,证明其具备通用性与实用性。

原文摘要 · Abstract (English)

Recent advancements in AI for biological research focus on integrating molecular data with natural language to accelerate drug discovery. However, the scarcity of high-quality annotations limits progress in this area. This paper introduces LA$^3$, a Language-based Automatic Annotation Augmentation framework that leverages large language models to augment existing datasets, thereby improving AI training. We demonstrate the effectiveness of LA$^3$ by creating an enhanced dataset, LaChEBI-20, where we systematically rewrite the annotations of molecules from an established dataset. These rewritten annotations preserve essential molecular information while providing more varied sentence structures and vocabulary. Using LaChEBI-20, we train LaMolT5 based on a benchmark architecture to learn the mapping between molecular representations and augmented annotations. Experimental results on text-based *de novo* molecule generation and molecule captioning demonstrate that LaMolT5 outperforms state-of-the-art models. Notably, incorporating LA$^3$ leads to improvements of up to 301% over the benchmark architecture. Furthermore, we validate the effectiveness of LA$^3$ notable applications in *image*, *text* and *graph* tasks, affirming its versatility and utility.

分子生成大模型数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。