用结构因果模型提升长链推理的准确性与可解释性。
Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

- 构建显式依赖图,将验证过程视为结构化推理而非盲目推演。
- 发现推理链长度与准确率呈倒U型关系,过长会降低性能。
- 通过分组相对策略优化动态平衡深度与简洁性,适合复杂事实验证场景。
多跳事实验证需在分散证据间进行复杂推理,对大语言模型构成挑战,易出现幻觉和逻辑断裂。现有方法虽通过思维链提升透明度,但缺乏对证据与论断间结构依赖的显式建模。本文提出受结构因果模型启发的框架,将推理过程锚定在显式的有向依赖图上,将验证视为构造性结构推理,而非完整因果推断。实证发现推理链长度与准确率呈倒U型相关,过度结构复杂性会损害性能。为此,我们提出基于规则的强化学习策略——分组相对策略优化(GRPO),动态优化结构深度与简洁性的权衡。在HoVer和EX-FEVER数据集上的大量实验表明,SCM-GRPO框架优于强基线,并生成更可追溯的复杂推理结构。
原文摘要 · Abstract (English)
Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models , which may suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought , often lack explicit modeling of the structural dependencies between evidence and claims. In this work, we introduce an SCM-inspired framework that grounds reasoning in explicit directed dependency graphs, treating verification as a constructive structural reasoning process rather than full causal inference with interventions or counterfactual semantics. We empirically identify an "inverted U-shaped" correlation between reasoning-chain length and accuracy, revealing that excessive structural complexity can degrade performance. To address this, we propose a rule-based reinforcement learning strategy using Group Relative Policy Optimization. This approach dynamically optimizes the trade-off between structural depth and conciseness. Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework outperforms strong baselines while producing more traceable reasoning structures for complex fact verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。