用结构化数据自动构建可扩展推理环境,提升大模型泛化能力。
Structured In-context Environment Scaling for Large Language Model Reasoning
- 基于大规模结构化数据自动生成推理环境,实现可扩展性
- 在域内任务上显著提升推理表现,并有效迁移至数学逻辑任务
- 支持信息缺失下的环境探索,增强模型鲁棒性与泛化能力
大语言模型通过强化学习和环境探索在推理能力上取得显著进展。环境的内在特性决定了模型能学习的能力,因此环境在强化学习微调中至关重要。理想的推理环境应具备可扩展性、可泛化的推理能力以及可验证性。然而,现有数学与编码环境因依赖专家标注而难以扩展,游戏类环境中的技能又过于专用,难以泛化。为此,我们提出结构化上下文环境(SIE)框架。SIE通过从大规模结构化数据自动构建推理环境实现可扩展性,其丰富的组合模式天然支持可泛化的推理。同时,结构化数据中的显式模式与推理链为规则化验证提供基础。实验表明,SIE不仅在域内结构化推理任务上取得显著提升,还使学到的组合推理能力有效迁移到域外的数学与逻辑推理任务。我们进一步探索了信息受限的局部SIE学习,发现模型可通过环境探索推断缺失信息,从而实现更强的推理性能与泛化能力。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental exploration. As the intrinsic properties of the environment determine the abilities that LLMs can learn, the environment plays a important role in the RL finetuning process. An ideal LLM reasoning environment should possess three core characteristics: scalability, generalizable reasoning, and verifiability. However, existing mathematical and coding environments are difficult to scale due to heavy reliance on expert annotation, while the skills learned in game-based environments are too specialized to generalize. To bridge this gap, we introduce the \textbf{S}tructured \textbf{I}n-context \textbf{E}nvironment (SIE) framework. SIE achieves scalability by automatically constructing reasoning environments from large-scale structured data, where the rich compositional patterns naturally support generalizable reasoning. Moreover, the explicit schemas and reasoning chains in structured data provide a foundation for rule-based verifiability. Experimental results show that SIE framework not only achieves substantial improvements in in-domain structured reasoning, but also enables the learned compositional reasoning skills to generalize effectively to out-of-domain mathematical and logical reasoning tasks. We further explored learning in information-limited partial SIEs and found that LLMs can infer the missing information through exploring the environment, leading to robust reasoning improvements and generalization performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。