arXiv:2603.20538cs.LGstat.ML2026-03

解释行为克隆中动作量化的理论原理与误差传播机制。

Understanding Behavior Cloning with Action Quantization

  • 分析动作量化的误差如何随时间累积并影响学习效果。
  • 证明在稳定动态下,量化行为克隆可达到最优样本复杂度。
  • 提出模型增强方法,在无需平滑策略假设下提升误差上限。

行为克隆是机器学习中的基础范式,广泛应用于机器人、自动驾驶和生成模型等领域。自回归模型(如Transformer)在大语言模型和视觉-语言-动作系统中表现卓越。然而,将此类模型用于连续控制时需对动作进行量化,这一普遍做法却缺乏理论理解。本文提供了该实践的理论基础:分析量化误差沿时间轴的传播及其与统计样本复杂度的交互。我们证明,采用量化动作与对数损失的行为克隆可实现最优样本复杂度,匹配现有下界,且量化误差仅随时间呈多项式依赖,前提是系统动态稳定且策略满足概率平滑性条件。我们进一步刻画不同量化方案是否满足这些条件,并提出一种基于模型的增强方法,可严格改进误差界而无需平滑性假设。最后,我们建立了同时涵盖量化误差与统计复杂性的根本限制。

原文摘要 · Abstract (English)

Behavior cloning is a fundamental paradigm in machine learning, enabling policy learning from expert demonstrations across robotics, autonomous driving, and generative models. Autoregressive models like transformer have proven remarkably effective, from large language models (LLMs) to vision-language-action systems (VLAs). However, applying autoregressive models to continuous control requires discretizing actions through quantization, a practice widely adopted yet poorly understood theoretically. This paper provides theoretical foundations for this practice. We analyze how quantization error propagates along the horizon and interacts with statistical sample complexity. We show that behavior cloning with quantized actions and log-loss achieves optimal sample complexity, matching existing lower bounds, and incurs only polynomial horizon dependence on quantization error, provided the dynamics are stable and the policy satisfies a probabilistic smoothness condition. We further characterize when different quantization schemes satisfy or violate these requirements, and propose a model-based augmentation that provably improves the error bound without requiring policy smoothness. Finally, we establish fundamental limits that jointly capture the effects of quantization error and statistical complexity.

行为克隆动作量化理论分析强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。