arXiv:2605.09968cs.LGmath.OC2026-05被引 1

提出统一框架,用可计算信号判断学习系统何时该停止。

Consolidation-Expansion Operator Mechanics:A Unified Framework for Adaptive Learning

  • 定义'顺序间隙'作为控制信号,衡量知识整合与拓展的顺序敏感性。
  • 顺序间隙变小并稳定时,系统已收敛,可安全停止迭代。
  • 适用于强化学习、持续学习等五类场景,尤其适合语言模型递归生成。

每个自适应学习系统都需在巩固已有知识和拓展新证据之间交替。本文提出「整合-扩展算子机制」(OpMech),以精确刻画这一结构。核心是‘顺序间隙’\Ogap(θ; e),表示整合算子$Q$与拓展算子$P_e$在给定知识状态下的非对易程度。由于顺序间隙可从系统自身轨迹计算,它成为实时控制信号:值大说明结果仍依赖操作顺序;当顺序间隙下降并保持较小,则后续处理难以改变结果。三个理论结果赋予该信号明确意义:顺序间隙沿收敛路径衰减;持续大的顺序间隙意味着系统尚未进入稳定态;基于顺序间隙的停止规则在无噪声与有界噪声环境下均有可证明保证。该框架适用于五类任务:多臂赌博机、强化学习、随机优化、持续学习和递归语言模型。我们在三个典型情形中给出顺序间隙可靠追踪收敛的条件,并详细开发语言模型应用,展示如何以证据驱动方式替代启发式停止规则与固定递归预算。

原文摘要 · Abstract (English)

Every adaptive learning system must alternate between two operations: consolidating what it already knows and expanding into new evidence. We propose \emph{Consolidation-Expansion Operator Mechanics} (OpMech), a framework that makes this structure precise. The central object is the \emph{order-gap} $\Ogap(θ; e)$, the degree to which a consolidation operator~$Q$ and an expansion operator~$P_e$ fail to commute at a given knowledge state. Because the order-gap is computable from the system's own trajectory, it serves as a real-time control signal: large values indicate that the system is still sensitive to the ordering of consolidation and expansion; once the order-gap falls and stays small, further processing is unlikely to change the outcome. Three results give the signal precise meaning: the order-gap decays along convergent trajectories; a persistently large order-gap implies the system is far from its settled state; and an order-gap-based stopping rule terminates with provable guarantees in both noiseless and bounded-noise settings. The framework applies across five domains: bandits, reinforcement learning, stochastic optimization, continual learning, and recursive language models. We give conditions under which the order-gap reliably tracks convergence in three representative cases. We develop the recursive language model application in detail, showing how OpMech replaces heuristic stopping rules and fixed recursion budgets with principled, evidence-driven alternatives.

自适应学习持续学习语言模型控制信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。