arXiv:2606.03057cs.LGcs.AI2026-06

对比九种分子表示法,发现不同任务需用不同表示,无万能方案。

Rethinking Molecular Text Representations for LLMs: An Empirical Study

论文配图:Rethinking Molecular Text Representations for LLMs: An Empirical Study
图 1 · 摘自论文原文
  • 系统测试九种分子表示在八项化学任务中的表现
  • IUPAC在分子检索中胜出,结构表示法在结构任务占优
  • 专用模型对SMILES依赖强,但泛化能力差,提示需任务适配

大型语言模型(LLMs)在分子任务中应用日益广泛,但何种分子表示更优仍不明确。我们开展系统性基准测试,评估九种分子表示在八个化学任务中对16个LLMs(五类模型家族)的表现。结果表明性能强烈依赖于表示方式,无单一表示在所有任务中领先:CML最佳,其次为MolJSON、InChI和标准SMILES。显式结构化文本表示(如CML、MolJSON)在结构任务中占优;IUPAC在语义任务中胜出,对全部16个模型均实现最优分子检索;而尽管SMILES在预训练中广泛使用,其在多数任务中并非最优。专用化学模型在使用SMILES时表现良好,但在结构化表示上显著退化,表明仅以SMILES评估会奖励不具泛化性的专精。通过LLM作为评判者,发现IUPAC生成正确分子的比例最高。通过分词审计、线性探针和注意力分析揭示,不同表示在模型内部编码机制不同:结构化表示需要更高的跨分子跨度注意力。研究反对表示无关的评估,提倡基于任务的表示路由策略。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use. We present a systematic benchmark evaluating LLM molecular competence across nine representations and eight chemical tasks. We benchmark 16 LLMs across five model families, including reasoning and non-reasoning variants, chemistry-specialized LLMs, and closed frontier models. Performance is strongly representation-dependent and no single representation wins across tasks, though CML is the best, followed by MolJSON, InChI, and then canonical SMILES. Explicit structured text representations (CML and MolJSON) dominate structural tasks; IUPAC dominates semantic tasks, winning molecule retrieval for all 16 LLMs; and SMILES variants are rarely optimal despite their prevalence in pretraining. Chemistry-specialized models perform well with SMILES at the cost of large degradations with structured text representations, suggesting SMILES-only evaluation rewards specialization that does not generalize. Using LLM-as-a-judge, we find that IUPAC produces the highest fraction of correct molecule generations. A mechanistic study via tokenization audits, linear probes and attention shows that representations are encoded differently inside the model; for example, structured representations require higher attention across the molecular span. Our results argue against representation-invariant evaluation and motivate task-aware representation routing for LLM-based chemistry.

分子表示大模型化学信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。