arXiv:2511.05933cs.CLcs.AI2025-11被引 2

强化学习让大模型更会找知识,而非记更多知识。

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

  • 用强化学习提升模型在参数内知识层级中的导航能力。
  • 在深层检索任务中,推理模型召回率随深度增长显著更高。
  • 适合关注模型知识检索效率的研究者和开发者。

强化学习(RL)常被认为以牺牲知识为代价提升语言模型推理能力。我们挑战这一观点,发现推理模型在纯知识回忆任务中始终优于指令微调版本。这些提升并非源于新知识,而是对模型内部已有知识层级的导航技能优化。结构化提示能有效恢复五种模型家族中指令与推理模型间的差距。在未见且不可提取的事实上的控制性强化学习实验中,模型对保留的高频但此前无法访问事实的召回率得到提升,排除了简单数据暴露的影响。在分层检索任务中,随着检索深度增加,推理模型表现出更优的层级遍历能力。逐层激活分析显示,尽管事实表征在指令与推理模型间保持高余弦相似度,查询表征却明显分化,表明推理主要改变的是知识遍历方式而非知识表示本身。最后发现,蒸馏模型常无法匹配推理模型的知识回忆表现,因其仅模仿自我修正而缺乏层级导航所需的探索行为。这些发现表明,提升大模型知识召回不仅需扩展知识量,还需训练其导航能力,推动未来后训练方法向优化知识遍历方向发展。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by showing that reasoning models consistently outperform their instruction-tuned versions on pure knowledge recall tasks. These gains do not reflect newly acquired information, but rather an improved procedural skill in navigating and searching existing knowledge hierarchies within the model parameters. Structured prompting, which explicitly guides models through hierarchical traversal -- recovers most of the instruct-reasoning gap across five model families. A controlled RL experiment on unseen, non-extractable facts improves recall of held-out frequent but previously inaccessible facts, ruling out simple data exposure. On depth-stratified retrieval tasks, reasoning models exhibit superior traversal as retrieval depth grows. Layerwise activation analysis further shows that while factual representations maintain high cosine similarity between instruct and reasoning models, query representations diverge noticeably, indicating that reasoning primarily reshapes how models traverse knowledge rather than the knowledge representation itself. Finally, we find that distilled models often fail to match reasoning models on knowledge recall because they imitate self-correction without acquiring the exploratory behavior needed for hierarchical navigation. Together, these findings suggest that improving factual recall in LLMs depends not only on expanding what models know but also on teaching them to navigate it -- motivating future post-training methods that optimize traversal.

强化学习知识导航推理能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。