arXiv:2606.15232cs.RO2026-06

提出PRISM模型,用多尺度空间注意力提升机器人抓取的视觉表征能力。

Rethinking Implicit Spatial Representation in Visuomotor Policy Learning

论文配图:Rethinking Implicit Spatial Representation in Visuomotor Policy Learning
图 1 · 摘自论文原文
  • 通过自上而下的注意力融合,保留多尺度空间信息。
  • 在低分辨率下成功率从5.0%提升至13.4%,参数仅增15.4%。
  • 适合高精度、低分辨率视觉输入的机器人操控任务。

基于生成模型的模仿学习已成为机器人操作的主流范式,策略性能高度依赖于条件化视觉表征。尽管先前方法采用空间Softmax池化,其有效性与机制仍不清晰。本研究系统比较不同池化方式,发现该操作可生成紧凑稳定的视觉特征,优于特征值表示,且维度更低。互补的显著性分析表明,此类表征使编码器更一致地关注任务相关区域。然而,现有视觉编码器中重复下采样导致细粒度空间信息衰减,限制了优势发挥。为此,我们提出PRISM,通过自上而下的交叉注意力融合,保留多尺度隐式空间信息。在多个任务和策略架构上实验均显示一致提升。尤其在低分辨率、高精度的ToolHang任务中,成功率从5.0%提升至13.4%,参数仅增加15.4%。结果支持多尺度隐式空间表征作为机器人操作的有效高效设计原则。

原文摘要 · Abstract (English)

Generative model-based imitation learning has become a widely adopted paradigm for robotic manipulation, where policy performance depends critically on the conditioned visual representations. Although spatial softmax-based representations have been adopted in prior visuomotor policies, their effectiveness and underlying mechanisms remain insufficiently understood. This work rethinks the use of spatial softmax pooling: do such implicit spatial representations provide effective and stable visual features for robotic manipulation? Through systematic studies of different pooling methods in visual encoders, we find that this pooling operation produces compact and stable spatial representations, which outperform feature-value representations, despite using substantially fewer dimensions. Complementary saliency analysis further suggests that these spatial representations guide the encoder to focus more consistently on task-relevant regions. However, this advantage is limited by a representation bottleneck in current visual encoders: repeated downsampling operations weaken fine-grained spatial information before the action-generation module can use it, especially under low-resolution observations. Motivated by these findings, we propose PRISM, a visual encoder that preserves multiscale implicit spatial information through top-down cross-attention fusion. Experiments across multiple tasks and policy backbones show consistent improvements. In particular, on the low-resolution, high-precision ToolHang task, PRISM shows clear gains, improving the average success rate from 5.0% to 13.4% while increasing parameters by only 15.4%. These results support the use of multiscale implicit spatial representations as an effective and efficient design principle for robotic manipulation.

机器人操控视觉表征空间注意力多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。