arXiv:2605.22493cs.LGcs.AI2026-05

解析动作分块行为克隆中的多模态失败机制

Understanding Multimodal Failure in Action-Chunking Behavioral Cloning

论文配图:Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
图 1 · 摘自论文原文
  • 对比潜变量与生成式策略的多模态建模差异
  • 过度正则化会丢失动作区分信息,正则不足则依赖先验覆盖范围
  • 生成策略需平滑映射或引入非支持区域以覆盖多模式

当同一观测对应多个有效动作时,行为克隆面临挑战。本文研究动作分块策略下的多模态问题,发现不同多模态参数化方式失效机制各异。对于潜变量策略,后验-先验正则化可提升部署时采样的可靠性,但过度正则化会消除动作条件信息,导致无法区分示范模式;适度降低正则化虽能保留模式信息,但成功率依赖先验是否覆盖相关潜空间区域。对于动作空间生成策略,多模态受限于基空间到动作空间映射的光滑性:小利普希茨常数的映射难以赋予多个分离模式高概率。因此,覆盖多模式需在基空间中引入尖锐转换或在动作空间中设置非支持桥接区域。合成多模态任务与机器人仿真基准实验验证了上述机制。

原文摘要 · Abstract (English)

Behavioral cloning becomes difficult when the same observation admits several valid actions. We study this problem for action-chunking policies and show that different multimodal parameterizations fail in different ways. For latent-variable policies, posterior-prior regularization makes deployment-time sampling more reliable, but excessive regularization removes the action-conditioned information needed to distinguish demonstrated modes. Reducing this regularization can preserve mode information, but then success depends on whether the prior covers the relevant latent regions. For action-space generative policies, multimodality is constrained by the smoothness of the base-to-action transport: a map with small Lipschitz constant cannot assign substantial probability to many well-separated modes. Covering many modes therefore requires either sharp transitions in base space or off-support bridge regions in action space. Experiments on synthetic multimodal tasks and robotic simulation benchmarks support these mechanisms.

行为克隆多模态动作分块

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。