首次完整刻画联合KL下长序列建模的误差规律,揭示其与序列长度的关系。
Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds
- 基于联合KL构建理论框架,分离近似与估计误差影响。
- 证明误差下界为Ω(H),与高效算法上界匹配,实现最优性证明。
- 适用于研究自回归模型理论性能的研究者,尤其关注序列生成与模仿学习。
我们研究在模型不匹配条件下,自回归建模与下一个词预测中长序列学习的根本性问题,使用联合Kullback--Leibler (KL) 散度进行度量。目标是厘清序列长度 $H$ 对联合分布、序列级误差中近似与估计误差的影响。通过建立匹配的上下界,我们首次在自然的联合KL目标下,完整刻画了长时域误差行为,相比现有工作提升了率并提供了最优性依据。近似方面,我们发现联合KL具有与序列长度无关的近似因子,与基于希尔伯特距离的分析形成鲜明对比(后者对计算高效的算法呈现Ω(H)依赖),凸显了散度选择对近似放大的决定作用。估计方面,我们证明了信息论下界为Ω(H),该结果对可分解策略类和完全共享策略均成立,与计算高效的算法所达到的$ ilde O(H)$上界一致。本分析通过精确的联合KLOracle理论,统一了日志损失训练目标、序列级评估指标与近似度量,进一步表明这些联合KL保证可导出与已有模仿学习文献相符的策略学习遗憾率。
原文摘要 · Abstract (English)
We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize how the sequence horizon \(H\) affects both approximation and estimation errors in this joint-distribution, sequence-level regime. By establishing matching upper and lower bounds, we provide, to our knowledge, the first complete characterization of long-horizon error behavior under the natural joint KL objective, with improved rates and optimality justification relative to existing work. On the approximation side, we show that joint KL admits a horizon-free approximation factor, in sharp contrast to Hellinger-based analyses that exhibit an \(Ω(H)\) dependence for computationally efficient methods; this isolates the choice of divergence as the source of approximation amplification. On the estimation side, we prove a fundamental information-theoretic lower bound of order \(Ω(H)\) that holds for both decomposable policy classes and fully shared policies, matching the \(\widetilde O(H)\) upper bounds achieved by computationally efficient algorithms. Our analysis clarifies the landscape of recent autoregressive learning results by aligning the log-loss training objective, the sequence-level evaluation metric, and the approximation metric {\color{black}through a sharp joint-KL oracle theory}. We further show that these joint-KL guarantees imply policy learning regret bounds at rates matching prior imitation learning literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。