探究大模型在材料科学中的知识可靠性,发现输出形式决定模型表现。
Probing Materials Knowledge in LLMs: From Latent Embeddings to Reliable Predictions
- 区分符号与数值任务,不同任务下微调策略效果差异显著。
- 直接提取中间层嵌入比输出文本预测更准,存在模型头瓶颈。
- 大模型性能随时间波动达9%-43%,影响科研可复现性。
大规模语言模型在材料科学中应用日益广泛,但其可靠性与知识编码机制仍不明确。我们在四个材料科学任务上评估了25个LLM,涵盖超过200种基础与微调配置。结果表明,输出模态从根本上决定了模型行为:符号任务中微调后答案趋于一致且可验证,响应熵降低;数值任务中微调虽提升预测精度,但重复推理结果不一致,限制其作为定量预测工具的可靠性。对于数值回归任务,从Transformer中间层提取嵌入比使用模型输出更有效,揭示了‘模型头瓶颈’现象,但该效应依赖于属性和数据集。最后,我们对GPT系列模型进行了18个月的纵向研究,发现其性能波动达9%-43%,给科学应用带来可复现性挑战。
原文摘要 · Abstract (English)
Large language models are increasingly applied to materials science, yet fundamental questions remain about their reliability and knowledge encoding. Evaluating 25 LLMs across four materials science tasks -- over 200 base and fine-tuned configurations -- we find that output modality fundamentally determines model behavior. For symbolic tasks, fine-tuning converges to consistent, verifiable answers with reduced response entropy, while for numerical tasks, fine-tuning improves prediction accuracy but models remain inconsistent across repeated inference runs, limiting their reliability as quantitative predictors. For numerical regression, we find that better performance can be obtained by extracting embeddings directly from intermediate transformer layers than from model text output, revealing an ``LLM head bottleneck,'' though this effect is property- and dataset-dependent. Finally, we present a longitudinal study of GPT model performance in materials science, tracking four models over 18 months and observing 9--43\% performance variation that poses reproducibility challenges for scientific applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。