构建量子材料研究的AI评估基准,推动智能科研发展
QMBench: A Research Level Benchmark for Quantum Materials Research
- 设计跨多领域的量子材料研究评测体系
- 涵盖结构、电子、热力学等多类科学问题
- 适合从事AI+材料科学的科研人员参考
我们提出QMBench,一个用于评估大语言模型代理在量子材料研究中能力的综合性基准。该基准聚焦于凝聚态物理知识与密度泛函理论等计算技术的应用,覆盖量子材料研究的多个方面,包括结构特性、电子性质、热力学及其他性质、对称性原理和计算方法。通过提供标准化的评估框架,QMBench旨在加速具备创造性贡献能力的AI科学家的发展。我们期待该基准能由研究社区持续改进与拓展。
原文摘要 · Abstract (English)
We introduce QMBench, a comprehensive benchmark designed to evaluate the capability of large language model agents in quantum materials research. This specialized benchmark assesses the model's ability to apply condensed matter physics knowledge and computational techniques such as density functional theory to solve research problems in quantum materials science. QMBench encompasses different domains of the quantum material research, including structural properties, electronic properties, thermodynamic and other properties, symmetry principle and computational methodologies. By providing a standardized evaluation framework, QMBench aims to accelerate the development of an AI scientist capable of making creative contributions to quantum materials research. We expect QMBench to be developed and constantly improved by the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。