arXiv:2605.21414cs.ROcs.CV2026-05中稿 · RSS 2026被引 4

用3D点云提升机器人动作精准度,让机械臂更懂空间细节。

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

论文配图:PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
图 1 · 摘自论文原文
  • 将多尺度点云融入动作生成,实现几何与语义协同决策。
  • 在RLBench-10Tasks上成功率提升10%,冻结视觉模型仍有效。
  • 适合需要高精度3D操作的机器人研究者和工程师。

视觉-语言-动作(VLA)模型通过利用预训练的视觉-语言骨干网络,在通用机器人操作中展现出强大潜力。然而,现有大多数VLA主要依赖2D视觉表征,难以处理细粒度几何与空间定位,限制了其在复杂3D环境中的精确操控能力。本文提出PointACT,一种双系统3D感知的VLA策略,直接将分层3D点云表示整合至动作解码过程。PointACT采用高效的瓶颈窗口自注意力机制,实现多尺度点-动作交互,使动态动作令牌能够同时关注局部几何细节与全局场景结构。我们在LIBERO和RLBench基准上评估PointACT,系统对比了单体与双系统基线模型,包括引入点云输入的变体。结果表明,PointACT在两个基准上均取得一致提升,在挑战性的RLBench-10Tasks套件上相比最先进预训练VLA成功率提高10%,且在冻结视觉语言骨干、仅从零训练动作专家时提升更显著。大量消融实验证明,将分层3D几何与预训练2D语义表征紧密耦合,是实现鲁棒、空间对齐机器人控制的关键。结果还凸显了预训练3D表征在3D感知VLA策略中的前景。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations, which limit their ability to reason about fine-grained geometry and spatial grounding - capabilities that are essential for precise and robust manipulation in 3D environments. In this paper, we propose PointACT, a dual-system 3D-aware VLA policy that integrates hierarchical 3D point cloud representations directly into the action decoding process. PointACT employs a multi-scale point-action interaction mechanism with efficient bottleneck window self-attention, enabling evolving action tokens to densely attend to both local geometric detail and global scene structure. We evaluate PointACT on the LIBERO and RLBench benchmarks and systematically compare it against monolithic and dual-system VLA baselines, including variants augmented with point cloud inputs. PointACT achieves consistent improvements across both benchmarks, increasing success rates by 10% on the challenging RLBench-10Tasks suite over state-of-the-art pretrained VLAs, with even larger gains when the vision-language backbone is frozen and the action expert is trained from scratch. Extensive ablation studies demonstrate that tightly coupling hierarchical 3D geometry with pretrained 2D semantic representations is critical for robust and spatially grounded robot control. Our results also highlight the promise of pretrained 3D representations for 3D-aware VLA policies.

机器人控制3D感知多模态点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。