arXiv:2605.01482cs.AI2026-05

用结构因果模型提升长链推理的准确性与可解释性。

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

论文配图:Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
图 1 · 摘自论文原文
  • 构建显式依赖图,将验证过程视为结构化推理而非盲目推演。
  • 发现推理链长度与准确率呈倒U型关系,过长会降低性能。
  • 通过分组相对策略优化动态平衡深度与简洁性,适合复杂事实验证场景。

多跳事实验证需在分散证据间进行复杂推理,对大语言模型构成挑战,易出现幻觉和逻辑断裂。现有方法虽通过思维链提升透明度,但缺乏对证据与论断间结构依赖的显式建模。本文提出受结构因果模型启发的框架,将推理过程锚定在显式的有向依赖图上,将验证视为构造性结构推理,而非完整因果推断。实证发现推理链长度与准确率呈倒U型相关,过度结构复杂性会损害性能。为此,我们提出基于规则的强化学习策略——分组相对策略优化(GRPO),动态优化结构深度与简洁性的权衡。在HoVer和EX-FEVER数据集上的大量实验表明,SCM-GRPO框架优于强基线,并生成更可追溯的复杂推理结构。

原文摘要 · Abstract (English)

Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models , which may suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought , often lack explicit modeling of the structural dependencies between evidence and claims. In this work, we introduce an SCM-inspired framework that grounds reasoning in explicit directed dependency graphs, treating verification as a constructive structural reasoning process rather than full causal inference with interventions or counterfactual semantics. We empirically identify an "inverted U-shaped" correlation between reasoning-chain length and accuracy, revealing that excessive structural complexity can degrade performance. To address this, we propose a rule-based reinforcement learning strategy using Group Relative Policy Optimization. This approach dynamically optimizes the trade-off between structural depth and conciseness. Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework outperforms strong baselines while producing more traceable reasoning structures for complex fact verification.

多跳推理因果模型强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。