arXiv:2606.01847cs.ROcs.LG2026-06中稿 · ed

纠正机器人动作策略中的几何错误,让模型更符合真实运动规律。

The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space

论文配图:The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space
图 1 · 摘自论文原文
  • 在SE(3)流形上直接进行扩散建模,避免传统方法的欧氏近似误差。
  • 任务长度提升7.3%,实机测试中多数任务表现优于基线。
  • 适合需要精准姿态控制的机器人操作场景,如灵巧抓取与装配。

基于扩散模型的视觉-语言-动作策略在机器人操作中取得显著成功,但存在根本性几何缺陷,即把SE(3)位姿近似为平坦的ℝ¹²向量,导致(1)流形漂移破坏SO(3)约束,(2)坐标变换下等变性失效,(3)轨迹非测地线且运动代价过高。本文提出Lie Diffuser Actor(LDA),一种在SE(3)上原生运行的扩散框架。通过左不变随机微分方程注入噪声,在切空间预测得分,并用指数映射重构采样结果。该方法从构造上消除流形漂移,保证坐标系等变性和测地线最优性。在CALVIN ABC→D数据集上,平均任务长度从3.27提升至3.51(+7.3%)。进一步在真实机器人上验证,方法在多数任务中优于基线。

原文摘要 · Abstract (English)

Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the $\textbf{Euclidean Fallacy}$: representing SE(3) poses as flat $\mathbb{R}^{12}$ vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce $\textbf{Lie Diffuser Actor (LDA)}$, a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABC$\rightarrow$D, LDA improves average task length from $3.27$ to $3.51$ ($+7.3\%$). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.

机器人扩散模型姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。