通过扩展词汇表提升化学语言模型表示能力
The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models
- 在预训练模型中加入化学相关词元,统一自然语言与分子结构表示
- 在化学文本上持续预训练后,下游任务性能显著提升
- 适合需要精准分子表示的药物发现与化学生成研究者
将大语言模型(LLM)应用于化学领域常受制于“分词瓶颈”:针对通用文本训练的分词器会将化学表示如SMILES拆分为语义无关的子词元。本文提出一种系统性方法,通过在预训练模型中针对性扩展化学相关词元,并在化学领域文本上继续预训练,实现自然语言与分子结构的统一表示。实验证明该策略能显著提升多种下游化学任务的性能。
原文摘要 · Abstract (English)
The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment chemical representations such as SMILES into semantically uninformative sub-tokens. This paper introduces a principled methodology to resolve this bottleneck by unifying the representation of natural language and molecular structures within a single model. Our approach involves targeted vocabulary extension-augmenting a pretrained LLM's vocabulary with chemically salient tokens, followed by continued pretraining on chemistry-domain text to integrate this new knowledge. We provide an empirical demonstration of the effectiveness of this strategy, showing that our methodology leads to superior performance on a range of downstream chemical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。