arXiv:2511.12036cs.CEcond-mat.mtrl-sci2025-11被引 1

用物理计算反馈训练语言模型,设计新型高温合金。

Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys

  • 用热力学相变计算生成奖励信号,优化语言模型。
  • 单个统一奖励信号让模型同时优化多种合金设计目标。
  • 首次实现基于物理原理的合金设计偏好学习,适合材料科研者。

我们将偏好学习应用于语言模型引导的新型结构合金设计任务。与以往聚焦生成稳定无机晶体的研究不同,本方法针对一种尚未充分探索的材料类别——体心立方/体心四方(BCC/B2)超合金,这类材料在极端环境中有潜在应用价值。采用三个开源大模型(LLaMA-3.1、Gemma-2 和 OLMo-2),我们证明了可通过单一统一的奖励信号,利用直接偏好优化(DPO)对语言模型进行多目标优化。该奖励信号来源于热力学相变计算,而非依赖成本高昂的人工或启发式反馈,具有科学依据。据我们所知,这是首次使用物理基础反馈对语言模型进行偏好调优以实现结构合金设计。该框架具备通用性与可扩展性,为物理科学领域智能设计空间探索提供了新路径。

原文摘要 · Abstract (English)

We apply preference learning to the task of language model-guided design of novel structural alloys. In contrast to prior work that focuses on generating stable inorganic crystals, our approach targets the synthesizeability of a specific structural class: BCC/B2 superalloys, an underexplored family of materials with potential applications in extreme environments. Using three open-weight models (LLaMA-3.1, Gemma-2, and OLMo-2), we demonstrate that language models can be optimized for multiple design objectives using a single, unified reward signal through Direct Preference Optimization (DPO). Unlike prior approaches that rely on heuristic or human-in-the-loop feedback (costly), our reward signal is derived from thermodynamic phase calculations, offering a scientifically grounded criterion for model tuning. To our knowledge, this is the first demonstration of preference-tuning a language model using physics-grounded feedback for structural alloy design. The resulting framework is general and extensible, providing a path forward for intelligent design-space exploration across a range of physical science domains.

合金设计偏好学习语言模型物理驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。