提出分层策略优化方法,实现多尺度稳定控制。
Fibration Policy Optimization
- 基于纤维丛结构分解强化学习数据,分离轨迹与词元级控制
- 新目标函数在策略在线时梯度为单位矩阵,提升更新效率
- 可扩展至四层架构,支持不同层级独立信任域预算
大型语言模型正日益以跨领域、专家划分和代理流水线等异构系统形式存在,但现有近端目标仅作用于单一尺度,缺乏对词元级、轨迹级及高层级分层稳定性控制的合理耦合机制。为此,我们推导出聚合策略裁剪目标(APC-Obj),首次实现基于样本的TV-TRPO的无约束精确重构,证明了裁剪型代理设计与信任域优化是同一问题的双重表述。在此基础上,我们提出纤维丛门控(FBG)——一种将采样强化学习数据组织为纤维丛的代数框架,将比例门控分解为基层面的轨迹聚合门控与纤维层面的词元残差门控,且在策略接近在线时具有一阶精确性。由此导出的纤维化策略优化(FiberPO)其雅可比矩阵在轨迹上块对角化,且在在线策略下退化为单位矩阵,提供更优更新方向,从而提升词元效率。该框架的组合性质超越轨迹-词元情形:纤维丛可代数组合成纤维门控层次结构(FGH),以相同门控机制扩展至任意深度,无需新增原语;如所提出的FiberPO-Domain,实现了包含领域、提示组、轨迹、词元四层级的独立信任域预算配置。这些成果将信任域理论、组合代数结构与实际多尺度稳定性控制统一为大型语言模型策略优化的完整框架。
原文摘要 · Abstract (English)
Large language models are increasingly trained as heterogeneous systems spanning multiple domains, expert partitions, and agentic pipelines, yet prevalent proximal objectives operate at a single scale and lack a principled mechanism for coupling token-level, trajectory-level, and higher-level hierarchical stability control. To bridge this gap, we derive the Aggregational Policy Censoring Objective (APC-Obj), the first exact unconstrained reformulation of sample-based TV-TRPO, establishing that clipping-based surrogate design and trust-region optimization are dual formulations of the same problem. Building on this foundation, we develop Fiber Bundle Gating (FBG), an algebraic framework that organizes sampled RL data as a fiber bundle and decomposes ratio gating into a base-level gate on trajectory aggregates and a fiber-level gate on per-token residuals, with provable first-order agreement with the true RL objective near on-policy. From APC-Obj and FBG we derive Fibration Policy Optimization (or simply, FiberPO), a concrete objective whose Jacobian is block-diagonal over trajectories, reduces to identity at on-policy, and provides better update direction thus improving token efficiency. The compositional nature of the framework extends beyond the trajectory-token case: fibrations compose algebraically into a Fibration Gating Hierarchy (FGH) that scales the same gating mechanism to arbitrary hierarchical depth without new primitives, as demonstrated by FiberPO-Domain, a four-level instantiation with independent trust-region budgets at the domain, prompt group, trajectory, and token levels. Together, these results connect the trust-region theory, a compositional algebraic structure, and practical multi-scale stability control into a unified framework for LLM policy optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。