arXiv:2604.07981cs.CLcs.AI2026-04被引 1

将长文本推理拆解为基本技能,提升大模型的复杂推理能力。

A Decomposition Perspective to Long-context Reasoning for LLMs

论文配图:A Decomposition Perspective to Long-context Reasoning for LLMs
图 1 · 摘自论文原文
  • 把长文本推理分解为可训练的基础技能
  • 在多个基准上平均提升7.7%(46.3%→54.0%)
  • 适合想提升模型长程推理能力的研究者

长文本推理对真实复杂应用至关重要,但仍是大语言模型的重大挑战。现有研究常忽视该任务的内在复杂性。本文突破整体视角,将长文本推理分解为一系列基础原子技能,并自动生成针对每项技能的伪数据集。实证分析表明,这些原子技能的掌握程度与整体长文本推理表现强相关。基于此,我们在伪数据集上使用强化学习优化模型的原子技能,以提升其通用长文本推理能力。跨Loogle、Loong、LongBench-v2、BrowscompLong、Ruler-qa2和MRCR等多个基准的实验表明,该方法平均性能提升7.7%(从46.3%提升至54.0%),显著优于强基线。

原文摘要 · Abstract (English)

Long-context reasoning is essential for complex real-world applications, yet remains a significant challenge for Large Language Models (LLMs). Despite the rapid evolution in long-context reasoning, current research often overlooks the internal complexity of the long-context reasoning task itself. In this paper, we move beyond this holistic view and decompose long-context reasoning into a set of fundamental atomic skills, and we then automatically synthesize a suite of pseudo datasets, each explicitly targeting a specific atomic skill. Our empirical analysis confirms that proficiency in these atomic skills is strongly correlated with general long-text reasoning performance. Building on this insight, we employ reinforcement learning on these pseudo datasets to sharpen the model's atomic skills, in the hope of boosting its general long-context reasoning ability. Extensive experiments across multiple benchmarks demonstrate the effectiveness of our approach: it outperforms a strong baseline by an average margin of 7.7\% (improving from 46.3\% to 54.0\%) across Loogle, Loong, LongBench-v2, BrowscompLong, Ruler-qa2, and MRCR.

长文本推理强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。