QwenLong-L1.5通过三步创新,显著提升超长文本推理与记忆管理能力。
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
- 构建多跳推理数据生成流水线,合成高难度长程问答任务。
- 提出自适应熵控优化算法,稳定长上下文强化学习训练过程。
- 设计记忆增强架构,支持百万级token以上任务的分步推理。
我们提出QwenLong-L1.5,通过系统性后训练创新,实现卓越的长上下文推理与记忆管理能力。核心技术突破包括:(1) 长上下文数据合成流水线:将文档拆解为原子事实及其关系,程序化生成可验证的多跳推理问题,规模化构建高质量训练数据,超越简单检索任务,推动真实长程推理能力发展;(2) 稳定的长上下文强化学习:引入任务平衡采样与任务特定优势估计,缓解奖励偏差,并提出自适应熵控策略优化(AEPO),动态调节探索与利用权衡;(3) 超长上下文的记忆增强架构:针对超长序列无法全量容纳的问题,设计多阶段融合强化学习训练框架,实现单次推理与迭代记忆处理的无缝结合。基于Qwen3-30B-A3B-Thinking,QwenLong-L1.5在长上下文推理基准上表现媲美GPT-5和Gemini-2.5-Pro,平均超越基线9.90分。在1M~4M token的超长任务中,其记忆代理框架相较基线提升9.48分。该能力还可迁移至科学推理、记忆工具使用及长对话等通用场景。
原文摘要 · Abstract (English)
We introduce QwenLong-L1.5, a model that achieves superior long-context reasoning capabilities through systematic post-training innovations. The key technical breakthroughs of QwenLong-L1.5 are as follows: (1) Long-Context Data Synthesis Pipeline: We develop a systematic synthesis framework that generates challenging reasoning tasks requiring multi-hop grounding over globally distributed evidence. By deconstructing documents into atomic facts and their underlying relationships, and then programmatically composing verifiable reasoning questions, our approach creates high-quality training data at scale, moving substantially beyond simple retrieval tasks to enable genuine long-range reasoning capabilities. (2) Stabilized Reinforcement Learning for Long-Context Training: To overcome the critical instability in long-context RL, we introduce task-balanced sampling with task-specific advantage estimation to mitigate reward bias, and propose Adaptive Entropy-Controlled Policy Optimization (AEPO) that dynamically regulates exploration-exploitation trade-offs. (3) Memory-Augmented Architecture for Ultra-Long Contexts: Recognizing that even extended context windows cannot accommodate arbitrarily long sequences, we develop a memory management framework with multi-stage fusion RL training that seamlessly integrates single-pass reasoning with iterative memory-based processing for tasks exceeding 4M tokens. Based on Qwen3-30B-A3B-Thinking, QwenLong-L1.5 achieves performance comparable to GPT-5 and Gemini-2.5-Pro on long-context reasoning benchmarks, surpassing its baseline by 9.90 points on average. On ultra-long tasks (1M~4M tokens), QwenLong-L1.5's memory-agent framework yields a 9.48-point gain over the agent baseline. Additionally, the acquired long-context reasoning ability translates to enhanced performance in general domains like scientific reasoning, memory tool using, and extended dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。