让机器人动作适应不同速度和姿势,提升泛化能力。
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

- 用时空解耦方法构建动作流形,分离路径与时间变化
- 在多个环境中实现更优的迁移性能,对速度变化更鲁棒
- 适合需要灵活适应动作速度和姿态的任务场景
在具身智能中,从有限数据中实现稳健泛化是核心挑战。现有方法通过回归绝对坐标,违背了广义协变性原则,导致策略绑定特定运动风格和固定速度。为此,我们提出广义动作流形(GAM)框架,通过结构解耦实现广义协变性。具体而言,GAM在两个正交维度上强制不变性:(1) 时间不变性,使用弧长参数化器将空间路径几何与时间动态解耦,确保对速度变化的鲁棒性;(2) 几何不变性,通过模式-仿射-分解机制将轨迹映射到姿态归一化坐标系中的标准“世界线”,区分不变几何模式与仿射调制,保障空间泛化能力。将GAM集成于结构化视觉-语言-动作(VLA)架构中,使稀疏示范可密集填充连续有效的动作流形。实证结果表明,GAM显著优于无几何感知基线,在迁移性和鲁棒性方面表现更优。
原文摘要 · Abstract (English)
Achieving robust generalization from limited data is a central challenge in embodied intelligence. Prevailing methods fail by regressing absolute coordinates, which violates the principle of general covariance. Fundamentally, this conflates the intrinsic task geometry with rigid execution patterns, binding policies to specific motion styles and fixed speeds. To resolve this, we propose the Generalized Action Manifold (GAM) framework that enforces general covariance through structural disentanglement. Specifically, GAM realizes the manifold by enforcing invariance across two orthogonal dimensions: (1) Temporal Invariance, utilizing an Arc-Length Parameterizer to orthogonalize the spatial path geometry from temporal dynamics, ensuring robustness to velocity variations; (2) Geometric Invariance, where a Schema-Affine-Factorization mechanism maps trajectories to canonical ``world lines'' in a pose-normalized coordinate frame. This distinguishes invariant geometric schemas from affine modulations, ensuring spatial generalizability. By integrating GAM within a structured Vision-Language-Action (VLA) architecture, we enable sparse demonstrations to densely populate a continuous, valid action manifold. Empirical results demonstrate that GAM enables superior transfer and robustness capabilities, outperforming geometry-agnostic baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。