arXiv:2605.29555cs.CL2026-05

用专家规则生成对比评价,让大模型学会有依据地选材料。

From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals

论文配图:From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals
图 1 · 摘自论文原文
  • 用专家规则和盲猜配对生成偏好信号,训练模型判断材料
  • 小模型在无外部检索下达到接近闭源模型的评估准确率
  • 适合自动化材料发现中需要低成本可靠评估的场景

随着候选生成和高通量实验的发展,材料发现的瓶颈已从性质预测转向大规模候选集的可靠评估。本文提出MaterEval框架,为同一候选材料自动生成两种评价:遵循专家规则并提供支持证据的知情判断,以及移除规则后的盲猜。将两者作为偏好数据,引导原本缺乏材料领域标准的通用大模型,从直觉判断转向基于显式证据的可靠评估。为平衡效率、成本与可靠性,进一步设计快慢推理机制,将大规模快速筛选与小样本深度审查分离。以高熵合金评估为例,结果显示,仅依赖内部知识的小型开源大模型在准确率、结论一致性与证据辨别力上均有显著提升,逼近规则驱动的闭源大模型性能。结果表明,专家规则可系统转化为可学习的偏好信号,实现低成本、可部署的材料评估模块,支撑自主材料发现闭环。

原文摘要 · Abstract (English)

As candidate generation and high-throughput experimentation advance, the primary bottleneck in materials discovery is shifting from property prediction to making reliable evaluations among massive candidate sets. We propose a Knowledge-Augmented Preference Signals Framework, MaterEval, that automatically produces, for the same candidate, two evaluations: an informed judgment that follows expert rules and provides supporting evidence, and a rule-removed blind guess. By pairing the two evaluations as preference data, we guide general-purpose large language models (LLMs), originally lacking materials-specific criteria, from intuitive judgment toward reliable evaluation supported by explicit evidence. To balance throughput, cost, and reliability, we further introduce a fast-slow reasoning scheme that decouples large-scale rapid screening from in-depth review on a small subset. Using high-entropy alloy (HEA) assessment as a case study, we show that, without external retrieval and relying solely on internalized capabilities, small open-source LLMs achieve substantial gains in accuracy, conclusion consistency, and evidence discrimination, approaching the performance of rule-based closed-source LLMs. These results demonstrate that expert rules can be systematically transformed into learnable preference signals, enabling a low-cost and deployable evaluation module for autonomous materials discovery loops.

材料发现大模型评估偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。