用强化学习规划问题,让大模型更准地查证信息真伪。
DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration

- 把复杂断言拆成问题集,通过知识图谱逐步探索证据。
- 在三个数据集上最高提升15%准确率,达84.57分。
- 适合需要透明可解释的假信息检测场景。
自动化事实核查对大语言模型仍是挑战,源于传统检索系统存在“查询脆弱性”。我们提出DeLIVeR(基于分解学习的信息可信度识别框架),将证据检索视为一种强化策略探索任务。DeLIVeR利用规划型LLM将复杂断言分解为针对性问题集,用于在结构化知识图谱中进行高精度证据搜索。通过组相对策略优化(GRPO)与奖励机制,优化规划器策略,重点提升结构多样性与判断准确性。在LIAR、FEVER和PolitiFact数据集上的评估显示,采用Qwen2.5-7B模型时,DeLIVeR分别取得83.73、84.57、79.70的最高F1分数,相比HippoRAG2提升10%-15%。该方法通过强化问题规划策略,有效弥合多跳推理差距,提供可审计、透明的可验证假信息检测路径。
原文摘要 · Abstract (English)
Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. We propose DeLIVeR (Decomposed Learning for Information-grounded Veracity Recognition), a framework that treats evidence retrieval as a reinforced strategic exploration task. DeLIVeR utilizes a Planner LLM to decompose complex claims into targeted question sets, which are used to traverse structured Knowledge Graphs (KGs) for high-precision evidence. We optimize the Planner's policy using Group Relative Policy Optimization (GRPO) with a reward system prioritizing structural diversity and verdict accuracy. Our evaluation on LIAR, FEVER, and PolitiFact shows that DeLIVeR significantly outperforms state-of-the-art baselines. Using Qwen2.5-7B, our framework achieved peak F1-scores of 83.73, 84.57, and 79.70 respectively, representing a 10-15% improvement over HippoRAG2. By shifting to a reinforced question-planning strategy, DeLIVeR effectively bridges multi-hop reasoning gaps and provides an auditable, transparent path for verifiable misinformation detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。