通过几何对齐让视觉语言与动作空间兼容,提升机器人学习效果。
LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment

- 用李代数映射线性化动作流形,构建物理可加的固定长度表示。
- 分层离散化动作表示,使其局部度量近似各向同性,匹配语义空间。
- 适用于需要精准动作理解的机器人任务,尤其适合跨模态学习场景。
本文从格罗莫夫-沃瑟斯坦视角研究视觉-语言-动作(VLA)学习,旨在使动作表征的关系几何与视觉语言嵌入的语义几何相兼容。由于语义空间拓扑为线性各向同性,而机器人动作的物理流形是非欧几里得且各向异性的,两者度量结构不一致,直接回归无法成立。为此,我们提出LAST(李代数动作空间分词器),通过两阶段变换重建动作空间:(1) 全局拓扑线性化:利用李代数映射将动作流形线性化,将轨迹转换为固定长度、物理可加的表示;(2) 局部度量离散化:分层离散化表示为模式与白化残差,生成近似各向同性的局部坐标图,与语义度量统计对齐。通过在全局与局部层面解决结构不匹配问题,LAST显著提升了VLA模型的收敛速度与泛化能力。
原文摘要 · Abstract (English)
We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action representations compatible with the semantic geometry of VL embeddings. However, this alignment is non-trivial due to the mathematical heterogeneity between the domains: the semantic space of vision-language is topologically linear and isotropic, whereas the physical manifold of robotic action is non-Euclidean and anisotropic. Their disjoint metric structures render direct regression ill-posed. To resolve this incompatibility, we introduce LAST (Lie-algebraic Action Space Tokenizer), which reconstructs the action space to establish local metric compatibility with the VL modality via a two-stage transformation: (1) Global Topological Linearization: linearizing the action manifold via Lie-algebraic mapping, converting trajectories into a fixed-length, physically additive representation. (2) Local Metric Discretization: hierarchically discretizing the representation into schemas and whitened residuals, yielding approximately isotropic local charts that are statistically aligned with the semantic metric. By resolving the structural mismatch at both global and local levels, LAST enables VLA models with superior convergence and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。