arXiv:2605.17862cs.LGcs.AI2026-05被引 1

解决大模型强化学习中异步训练的性能损失问题

$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control

  • 用样本新鲜度评分动态调节过时数据影响
  • 在长序列任务中达到同步训练的性能水平
  • 适合需要高吞吐的智能体后训练场景

将大语言模型的在线策略蒸馏(OPD)扩展至长交互周期面临根本矛盾:为提升系统效率需异步执行,但会偏离理想的在线策略目标。本文理论分解目标偏差为轨迹漂移与监督漂移,分别反映学生模型轨迹和教师上下文的过时性。基于此提出样本级新鲜度评分,量化缓冲样本对在线目标的可靠性。据此设计f-OPD框架,自适应调控过时样本影响,约束异步训练下的策略漂移。在推理、工具使用和编码代理等交互周期递增的任务上,f-OPD性能接近同步优化,同时保持异步执行的吞吐优势。该工作首次实现OPD中性能与效率的平衡,为大规模长周期智能体后训练铺平道路。

原文摘要 · Abstract (English)

Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency, but structurally deviates from the ideal on-policy objective. To address this challenge, we theoretically decompose the objective discrepancy into rollout drift and supervision drift, capturing staleness in student rollout and teacher context, respectively. Building on this, we introduce a sample-level freshness score that quantifies the reliability of a buffered sample with respect to the on-policy objective. Guided by this signal, we further propose f-OPD, a novel framework that adaptively regulates stale-sample influence and constrains policy drift accumulated under asynchronous training. Across reasoning, tool-use, and coding-agent tasks of increasing interaction horizon, f-OPD consistently achieves task performance comparable to synchronous optimization while largely retaining the throughput advantages of asynchronous execution. Our results establish the first recipe for achieving a performance-efficiency trade-off in OPD, paving the way for long-horizon agentic post-training at scale.

强化学习大模型蒸馏异步训练智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。