arXiv:2602.05818cs.AIcs.DB2026-02被引 1

用智能体强化学习让大模型动态推理时序知识图谱,减少幻觉并提升泛化能力。

TKG-Thinker: Towards Dynamic Reasoning over Temporal Knowledge Graphs via Agentic Reinforcement Learning

  • 构建可自主规划与自适应检索的智能体,通过多轮交互进行时序推理。
  • 在三个开源大模型上测试,性能超越现有方法,复杂场景下泛化能力强。
  • 结合思维链微调与多维奖励强化学习,有效缓解时间约束下的推理幻觉。

时序知识图谱问答(TKGQA)旨在利用时序知识库回答依赖时间的问题。尽管大语言模型(LLMs)在TKGQA中展现出巨大潜力,但现有提示策略存在两大局限:一是复杂时间约束下易产生推理幻觉;二是静态提示限制了模型自主性与泛化能力,缺乏与时序知识图谱(TKGs)环境的动态交互优化。为此,我们提出新型智能体TKG-Thinker,具备自主规划与自适应检索能力,可通过双阶段训练实现深入时序推理。首先使用思维链数据进行监督微调(SFT),注入核心规划能力;随后通过强化学习(RL)阶段,利用多维奖励优化复杂时间约束下的推理策略。在基准数据集上的实验表明,TKG-Thinker在三种开源大模型上均达到领先性能,并展现出强大的跨复杂TKGQA场景泛化能力。

原文摘要 · Abstract (English)

Temporal knowledge graph question answering (TKGQA) aims to answer time-sensitive questions by leveraging temporal knowledge bases. While Large Language Models (LLMs) demonstrate significant potential in TKGQA, current prompting strategies constrain their efficacy in two primary ways. First, they are prone to reasoning hallucinations under complex temporal constraints. Second, static prompting limits model autonomy and generalization, as it lack optimization through dynamic interaction with temporal knowledge graphs (TKGs) environments. To address these limitations, we propose \textbf{TKG-Thinker}, a novel agent equipped with autonomous planning and adaptive retrieval capabilities for reasoning over TKGs. Specifically, TKG-Thinker performs in-depth temporal reasoning through dynamic multi-turn interactions with TKGs via a dual-training strategy. We first apply Supervised Fine-Tuning (SFT) with chain of thought data to instill core planning capabilities, followed by a Reinforcement Learning (RL) stage that leverages multi-dimensional rewards to refine reasoning policies under intricate temporal constraints. Experimental results on benchmark datasets with three open-source LLMs show that TKG-Thinker achieves state-of-the-art performance and exhibits strong generalization across complex TKGQA settings.

时序知识图谱智能体强化学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。