用可验证奖励训练模型,提升无机材料组合推理能力。
Mat-Pref: Verifiable-Reward Training Improves Compositional Reasoning in Inorganic Materials

- 基于密度泛函理论构建10837个离子替换问题,分三类评估泛化性。
- 微调+组相对策略优化使小模型在未见结构族上达71.6%准确率。
- 通过答案凝聚度分析揭示正确答案从可达成到成主流的机制突破。
可验证奖励强化学习(RLVR)在数学与代码推理中进展迅速,但扩展至科学领域时,现有基准无法分解泛化来源:性能提升是源于结构迁移、性质迁移,还是单纯记忆?我们提出Mat-Pref,一个包含10,837个离子替换问题的基准,覆盖11种无机结构家族,基于材料项目(Materials Project)的密度泛函理论计算构建。该基准包含三个评估分割:分布内表现、完全未见结构家族的泛化,以及跨性质转移——即仅通过生成能监督,将带隙推理应用于训练中见过的基质。四个零样本前沿模型(70-671B参数)在所有分割上均保持33%-54%准确率,表明规模本身不足以解决此类组合化学推理任务。采用两阶段流程:监督微调后接组相对策略优化(GRPO),使Qwen3-8B在分布内达到65.2%,在未见结构族上达71.6%,超越零样本的Qwen3-235B超过20个百分点。自一致性采样显示,微调后的策略虽可生成正确答案,却难以使其成为主导响应;而GRPO重构了响应分布,使正确答案成为主导而非仅可达。对关键决策层的逻辑透镜分析揭示,答案凝聚度提升约20个百分点。我们据此提出干扰项置换一致性指标,该指标下GRPO将宽松评分(至少一个置换正确)与严格评分(所有置换正确)的差距由24.0%缩小至14.3%。
原文摘要 · Abstract (English)
Reinforcement learning from verifiable rewards (RLVR) has driven rapid progress in mathematical and code reasoning, but when extended to science, existing benchmarks do not decompose what generalizes: do gains reflect structural transfer, property transfer, or memorization? We introduce Mat-Pref, a benchmark of 10,837 ionic-substitution questions across 11 inorganic structure families, grounded in density functional theory calculations from the Materials Project, with three evaluation splits that isolate in-distribution performance, generalization to entirely held-out structure families, and cross-property transfer: applying band-gap reasoning to hosts seen during training only through formation-energy supervision. Four zero-shot frontier models (70-671B parameters) remain in the 33-54% range on every split, confirming that scale alone does not resolve the compositional chemical reasoning this task demands. A two-stage pipeline of supervised fine-tuning followed by Group Relative Policy Optimization (GRPO) lifts Qwen3-8B to 65.2% in-distribution and 71.6% on held-out families, exceeding zero-shot Qwen3-235B by over 20 percentage points on both structural-generalization splits. Self-consistency sampling shows that the SFT policy can already produce correct answers but cannot reliably surface them as the modal response; GRPO reshapes the distribution so that correct answers become modal rather than merely reachable, and this sharper commitment is visible mechanistically: logit lens analysis reveals a ${\sim}$20pp advantage in answer crystallization at the critical decision layer. We formalize this observation as a distractor-permutation consistency metric under which GRPO narrows the gap between lenient scoring (at least one permutation correct) and strict scoring (all permutations correct) from 24.0 to 14.3 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。