评测大模型在材料科学问答与性质预测中的表现与鲁棒性。
Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions
- 用三种数据集测试大模型在材料科学任务中的表现。
- 发现模型在噪声和对抗干扰下性能下降明显,存在模式崩溃现象。
- 适合关注AI可靠性或材料科研自动化的人参考。
大语言模型(LLMs)有潜力革新科学研究,但其在特定领域的稳健性和可靠性仍不明确。本研究评估了大模型在材料科学中的表现与鲁棒性,聚焦于领域内问答和材料性质预测,在多种真实与对抗性条件下进行测试。使用三个数据集:1)本科材料科学课程的多项选择题;2)包含不同钢成分与屈服强度的数据集;3)包含材料晶体结构文本描述与带隙值的带隙数据集。通过零样本思维链、专家提示和少样本上下文学习等多种提示策略评估模型表现。在真实扰动与故意对抗性篡改下测试模型鲁棒性,以评估其在现实场景中的可靠性。研究还揭示了预测任务中独特的现象,如提示示例邻近度变化引发的模式崩溃,以及训练/测试分布不一致时的性能恢复。结果旨在为大模型在材料科学中的广泛应用提供审慎认知,并推动其鲁棒性与可靠性的改进。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have the potential to revolutionize scientific research, yet their robustness and reliability in domain-specific applications remain insufficiently explored. In this study, we evaluate the performance and robustness of LLMs for materials science, focusing on domain-specific question answering and materials property prediction across diverse real-world and adversarial conditions. Three distinct datasets are used in this study: 1) a set of multiple-choice questions from undergraduate-level materials science courses, 2) a dataset including various steel compositions and yield strengths, and 3) a band gap dataset, containing textual descriptions of material crystal structures and band gap values. The performance of LLMs is assessed using various prompting strategies, including zero-shot chain-of-thought, expert prompting, and few-shot in-context learning. The robustness of these models is tested against various forms of 'noise', ranging from realistic disturbances to intentionally adversarial manipulations, to evaluate their resilience and reliability under real-world conditions. Additionally, the study showcases unique phenomena of LLMs during predictive tasks, such as mode collapse behavior when the proximity of prompt examples is altered and performance recovery from train/test mismatch. The findings aim to provide informed skepticism for the broad use of LLMs in materials science and to inspire advancements that enhance their robustness and reliability for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。