让大模型的错误描述反哺药物发现,意外提升分子性质预测准确率。
Can Hallucinations Help? Boosting LLMs for Drug Discovery
- 用分子SMILES生成自然语言描述,利用常有的幻觉文本增强预测。
- Falcon3-Mamba-7B在加入幻觉后表现最优,GPT-4o生成幻觉提升最显著。
- 1.8万条有益幻觉中结构误述最有效,大模型更受益于幻觉。
大语言模型(LLMs)中的幻觉——即看似合理但事实错误的文本——通常被视为缺陷。本文探讨幻觉是否能提升LLMs在分子性质预测这一早期药物发现关键任务中的表现。我们提示模型从分子SMILES字符串生成自然语言描述,并将这些常含幻觉的描述用于下游分类任务。在五个数据集上评估七种指令微调的LLM,结果表明部分模型在引入幻觉后预测准确率显著提升。值得注意的是,Falcon3-Mamba-7B在包含幻觉时超越所有基线,而由GPT-4o生成的幻觉带来的增益最大。我们进一步识别并分类了超过1.8万条有益幻觉,发现结构误述是最具影响力的类型,暗示关于分子结构的幻觉可能增强模型置信度。消融实验显示,更大模型更受益于幻觉,温度调节影响较小。研究挑战了幻觉纯为负面的传统认知,提示其在药物发现等科学建模任务中可作为有效信号加以利用。
原文摘要 · Abstract (English)
Hallucinations in large language models (LLMs), plausible but factually inaccurate text, are often viewed as undesirable. However, recent work suggests that such outputs may hold creative potential. In this paper, we investigate whether hallucinations can improve LLMs on molecule property prediction, a key task in early-stage drug discovery. We prompt LLMs to generate natural language descriptions from molecular SMILES strings and incorporate these often hallucinated descriptions into downstream classification tasks. Evaluating seven instruction-tuned LLMs across five datasets, we find that hallucinations significantly improve predictive accuracy for some models. Notably, Falcon3-Mamba-7B outperforms all baselines when hallucinated text is included, while hallucinations generated by GPT-4o consistently yield the greatest gains between models. We further identify and categorize over 18,000 beneficial hallucinations, with structural misdescriptions emerging as the most impactful type, suggesting that hallucinated statements about molecular structure may increase model confidence. Ablation studies show that larger models benefit more from hallucinations, while temperature has a limited effect. Our findings challenge conventional views of hallucination as purely problematic and suggest new directions for leveraging hallucinations as a useful signal in scientific modeling tasks like drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。