用4维高斯场建模动态场景,提升机器人操作的感知与动作精度。
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
- 引入高斯动作场(GAF),在3D高斯点云基础上增加运动属性,实现4维动态建模。
- 重建质量提升:PSNR增11.54dB,SSIM增0.39,LPIPS降0.56,任务成功率提高7.3%。
- 适合需要精准动态场景理解的机器人抓取与操作任务,尤其在复杂运动环境下表现优异。
精准的场景感知对视觉驱动的机器人操作至关重要。现有方法多采用视觉到动作(V-A)或视觉到3D再到动作(V-3D-A)范式,但难以应对操作场景的复杂性和动态性,导致动作不准。本文提出视觉-4维-动作(V-4D-A)框架,通过高斯动作场(GAF)实现从运动感知的4维表示直接推理动作。GAF在3D高斯点云基础上引入可学习的运动属性,支持动态场景与操作动作的4维建模。GAF输出三重信息:当前场景重建、未来帧预测、初始动作估计(基于高斯运动)。同时,采用动作-视觉对齐的去噪框架,以统一表示(初始动作+高斯感知)为条件,进一步优化动作精度。大量实验表明,相比最先进方法,GAF在重建质量上实现PSNR提升11.5385 dB,SSIM提升0.3864,LPIPS降低0.5574,机器人操作任务平均成功率提升7.3%。
原文摘要 · Abstract (English)
Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action (V-A) paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action (V-3D-A) paradigm, leveraging intermediate 3D representations. However, these methods often struggle with action inaccuracies due to the complexity and dynamic nature of manipulation scenes. In this paper, we adopt a V-4D-A framework that enables direct action reasoning from motion-aware 4D representations via a Gaussian Action Field (GAF). GAF extends 3D Gaussian Splatting (3DGS) by incorporating learnable motion attributes, allowing 4D modeling of dynamic scenes and manipulation actions. To learn time-varying scene geometry and action-aware robot motion, GAF provides three interrelated outputs: reconstruction of the current scene, prediction of future frames, and estimation of init action via Gaussian motion. Furthermore, we employ an action-vision-aligned denoising framework, conditioned on a unified representation that combines the init action and the Gaussian perception, both generated by the GAF, to further obtain more precise actions. Extensive experiments demonstrate significant improvements, with GAF achieving +11.5385 dB PSNR, +0.3864 SSIM and -0.5574 LPIPS improvements in reconstruction quality, while boosting the average +7.3% success rate in robotic manipulation tasks over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。