让大模型推理规范性问题时不仅答对,还懂为什么——用可验证的环境和新算法提升逻辑严谨性。
NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO

- 构建可验证的规则世界环境,自动生成答案、证明与反例证书
- 新方法MR-GRPO在关系结构上准确率提升至0.99,减少错误归类
- 适合研究模型逻辑推理、可信AI及训练评估的科研人员
语言模型可能因错误理由得出正确规范结论。我们提出NormWorlds-CF,一个用于可执行规则世界中反事实规范推理的可验证环境。其确定性求解器能生成最终答案、证明与证伪证书、论点状态、支持集及成对世界变化标签,实现无需大模型评判的监督与评估。基准包含分阶段SFT诊断与紧凑的成对世界任务,共270个根族与1080个标准-变体对。SFT诊断显示,仅以答案监督可达到完美答案准确率,但无法获得证伪能力:答案仅训练达完美准确率但联合证伪证书得分为零;全混合训练加针对性重播则实现强任务整体准确率(0.99)。针对结构化变化任务,引入类条件奖励的元关系梯度策略优化(MR-GRPO),对关系族和求解器可见变化字段给予部分奖励。在匹配的Qwen3-1.7B续写实验中,相较于稀疏与仅答案奖励,MR-GRPO提升保留关系准确率与关系族正确性,降低错误族错误率。在Qwen3-4B三种子验证中,稀疏奖励最佳保持粗粒度关系标签,仅答案奖励改善答案变化但削弱关系族结构,而MR-GRPO在答案、支持集、状态变化字段以及类条件元关系与变化存在性上均领先。结果表明,经验证的反事实结构可塑造后训练过程,但完整变化记录生成、不变子类型识别与分布外(OOD)迁移仍为开放问题。
原文摘要 · Abstract (English)
Language models can reach the right normative verdict for the wrong reason. We introduce NormWorlds-CF, a solver-verified environment for counterfactual normative reasoning in executable rule worlds. Its deterministic solver produces final answers, proof and falsification certificates, argument statuses, support sets, and paired-world change labels, enabling supervision and evaluation without LLM judges. The benchmark contains staged SFT diagnostics and a compact paired-world task with 270 root families and 1080 canonical-to-variant pairs. The SFT diagnostics show that final-answer supervision can saturate verdict accuracy without inducing falsification competence: answer-only SFT reaches perfect answer accuracy but scores zero on joint falsification certificates, while full-mix training with targeted replay reaches strong all-task accuracy (0.99). For the structured-change task, we introduce metamorphic-relation GRPO (MR-GRPO), a class-conditioned reward for GRPO that gives partial credit for relation families and solver-visible change fields. In matched Qwen3-1.7B continuation experiments, MR-GRPO improves held-out relation accuracy and relation-family correctness, and reduces wrong-family error, compared to sparse and answer-only GRPO. In Qwen3-4B three-seed validation, sparse reward preserves coarse relation labels best, answer-only reward improves answer-change but weakens relation-family structure, and MR-GRPO leads on answer-, support-, and status-change fields as well as class-conditioned MR and change-presence. These results show that verified counterfactual structure can shape post-training beyond final answers, while exact full change-record generation, invariant subtype recognition, and out-of-distribution (OOD) transfer remain open problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。