arXiv:2510.12971cs.RO2025-10被引 4

仅用2-3段视频,让机器人学会可迁移的6自由度抓取技能。

Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation

  • 通过神经可操作函数提取物体几何、外观与可用性特征
  • 13项任务平均成功率提升14.9个百分点,仅需2-3个示范视频
  • 适合需要快速学习新操作的机器人部署场景

我们提出Actron3D框架,仅需少量单目、未校准、仅含RGB的人类视频,即可让机器人掌握可迁移的6-DoF操作技能。核心是神经可操作函数,一种紧凑的物体中心表示,将多样化未校准视频中的几何、视觉外观和可操作性信息提炼为轻量级神经网络,形成操作技能的记忆库。部署时,采用从粗到精的优化流程,通过连续查询神经函数中编码的多模态特征,检索相关可操作函数并转移精确的6-DoF操作策略。仿真与真实世界实验均表明,Actron3D显著优于现有方法,在13项任务上平均成功率提升14.9个百分点,且每任务仅需2-3个示范视频。

原文摘要 · Abstract (English)

We present Actron3D, a framework that enables robots to acquire transferable 6-DoF manipulation skills from just a few monocular, uncalibrated, RGB-only human videos. At its core lies the Neural Affordance Function, a compact object-centric representation that distills actionable cues from diverse uncalibrated videos-geometry, visual appearance, and affordance-into a lightweight neural network, forming a memory bank of manipulation skills. During deployment, we adopt a pipeline that retrieves relevant affordance functions and transfers precise 6-DoF manipulation policies via coarse-to-fine optimization, enabled by continuous queries to the multimodal features encoded in the neural functions. Experiments in both simulation and the real world demonstrate that Actron3D significantly outperforms prior methods, achieving a 14.9 percentage point improvement in average success rate across 13 tasks while requiring only 2-3 demonstration videos per task.

机器人操作视频学习可迁移技能6-DoF控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。