测试大模型是否真会分子属性预测,还是靠记忆答案。
In-Context Molecular Property Prediction with LLMs: A Blinding Study on Memorization and Knowledge Conflicts
- 用逐步隐藏信息的方法,检验模型是否依赖记忆。
- 在三个分子数据集上,未发现模型直接复制答案。
- 揭示了预训练知识与上下文信息的冲突,适合可信评估研究者。
大型语言模型(LLMs)的能力已从自然语言处理扩展到科学预测任务,包括分子属性预测。然而,其在上下文学习中的有效性仍不明确,尤其考虑到常用基准中存在训练数据污染的可能。本文研究了LLMs在分子属性上的上下文回归是否真实有效,或仅依赖对目标值的逐字记忆。通过一系列逐步盲化实验,分析预训练知识与上下文信息之间的相互作用。我们在三个分子数据集(Delaney溶解度、Lipophilicity、QM7原子化能)上,对九种不同家族的模型(GPT-4.1、GPT-5、Gemini 2.5)进行评估,采用系统性盲化策略逐步减少可用信息,并辅以0、60、1000次上下文样本量作为信息获取的控制变量。为验证记忆分析和盲化实验,引入正负控制及结构参考基线,并为所有结果提供自助法置信区间。结果显示,在传统基准上未发现逐字检索证据,且盲化暴露了预训练知识与上下文信息间的冲突。本工作为在受控信息访问下评估分子属性预测提供了原则性框架。
原文摘要 · Abstract (English)
The capabilities of large language models (LLMs) have expanded beyond natural language processing to scientific prediction tasks, including molecular property prediction. However, their effectiveness in in-context learning remains ambiguous, particularly given the potential for training data contamination in widely used benchmarks. This paper investigates whether LLMs perform genuine in-context regression on molecular properties or instead rely on verbatim retrieval of memorized target values. Furthermore, we analyze the interplay between pre-trained knowledge and in-context information through a series of progressively blinded experiments. We evaluate nine LLM variants across three families (GPT-4.1, GPT-5, Gemini 2.5) on three MoleculeNet datasets (Delaney solubility, Lipophilicity, QM7 atomization energy) using a systematic blinding approach that iteratively reduces available information, complemented by 0-, 60-, and 1000-shot in-context sample sizes as an additional control for information access. To validate the memorization analysis and the blinding experiments, we add a positive and a negative control for the memorization experiments and structural reference baselines for the multi-shot experiments as well as bootstrap confidence intervals for all results. We find no evidence of verbatim retrieval on the legacy benchmarks and show that blinding exposes conflicts between pre-trained knowledge and in-context information. This work provides a principled framework for evaluating molecular property prediction under controlled information access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。