长任务跨度会引发训练不稳,缩短任务可提升模型表现与泛化能力。
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length

- 通过控制任务规则,仅改变动作序列长度来研究其影响。
- 任务越长越难训练,易出现探索困难和奖励分配失准。
- 短任务训练的模型能更好适应更长任务,具备跨跨度泛化能力。
大型语言模型(LLMs)在通过多步交互解决复杂任务方面展现出潜力。然而,以往研究多关注系统优化或算法改进,对任务跨度长度如何影响训练过程的理解仍不充分。本文通过构建受控任务,使智能体面对相同的决策规则和推理结构,仅在完成任务所需的动作序列长度上不同。结果表明,单纯增加任务跨度即构成训练瓶颈,导致严重训练不稳定,根源在于探索困难和信用分配难题。我们证明,降低任务跨度是缓解该问题的关键策略,可稳定训练并提升长跨度任务表现。此外,发现低跨度训练的模型在推理时对更长跨度任务具有更强泛化能力,这一现象称为‘跨度泛化’。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promise as interactive agents that solve tasks through extended sequences of environment interactions. While prior work has primarily focused on system-level optimizations or algorithmic improvements, the role of task horizon length in shaping training dynamics remains poorly understood. In this work, we present a systematic empirical study that examines horizon length through controlled task constructions. Specifically, we construct controlled tasks in which agents face identical decision rules and reasoning structures, but differ only in the length of action sequences required for successful completion. Our results reveal that increasing horizon length alone constitutes a training bottleneck, inducing severe training instability driven by exploration difficulties and credit assignment challenges. We demonstrate that horizon reduction is a key principle to address this limitation, stabilizing training and achieving better performance in long-horizon tasks. Moreover, we find that horizon reduction is related to stronger generalization across horizon lengths: models trained under reduced horizons generalize more effectively to longer-horizon variants at inference time, a phenomenon we refer to as horizon generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。