arXiv:2512.00074cs.ROcs.CV2025-12中稿 · CVPR被引 2

让机器人视觉模型学会动态因果关系,提升操作成功率。

Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning

  • 将状态预测建模为生成扩散过程,联合学习正反向动态
  • 在16个仿真和4个真实任务中显著提升操作成功率
  • 无需动作或重建监督,适合大规模机器人学习场景

尽管在识别与分割任务上表现优异,现有3D视觉预训练方法在机器人操作任务中仍表现不足。我们归因于两点:缺乏状态-动作-状态动态建模,以及显式几何重建带来的冗余。本文提出AFRO,一种自监督框架,无需动作或重建标注即可学习动态感知的3D表征。AFRO将状态预测建模为生成扩散过程,并在共享潜在空间中联合建模前向与逆向动态,以捕捉因果转移结构。为防止动作学习中的特征泄露,采用特征差分与逆一致性监督,提升了视觉特征的质量与稳定性。结合Diffusion Policy,AFRO在16个仿真任务和4个真实世界任务中显著提高操作成功率,优于现有预训练方法。该框架随数据量和任务复杂度增长表现出良好可扩展性。定性可视化显示,AFRO学习到语义丰富且具有区分性的特征,为机器人3D表征学习提供有效预训练方案。

原文摘要 · Abstract (English)

Despite strong results on recognition and segmentation, current 3D visual pre-training methods often underperform on robotic manipulation. We attribute this gap to two factors: the lack of state-action-state dynamics modeling and the unnecessary redundancy of explicit geometric reconstruction. We introduce AFRO, a self-supervised framework that learns dynamics-aware 3D representations without action or reconstruction supervision. AFRO casts state prediction as a generative diffusion process and jointly models forward and inverse dynamics in a shared latent space to capture causal transition structure. To prevent feature leakage in action learning, we employ feature differencing and inverse-consistency supervision, improving the quality and stability of visual features. When combined with Diffusion Policy, AFRO substantially increases manipulation success rates across 16 simulated and 4 real-world tasks, outperforming existing pre-training approaches. The framework also scales favorably with data volume and task complexity. Qualitative visualizations indicate that AFRO learns semantically rich, discriminative features, offering an effective pre-training solution for 3D representation learning in robotics. Project page: https://kolakivy.github.io/AFRO/

机器人学习动态建模自监督3D表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。