用低分辨率视觉通过行为克隆实现机器人主动感知与抓取
Behavior Cloning for Active Perception with Low-Resolution Egocentric Vision
- 直接从低分辨率图像预测关节增量动作
- 低分辨率视觉下任务成功率高,相对位移预测更优
- 适合低成本机器人视觉系统与主动感知研究者
我们研究行为克隆是否足以在结构化物体寻找任务中实现主动感知。一个配备腕装全景RGB相机的低成本机械臂需重新定位,使部分可见的植物居中后触发抓取信号,这要求动作能改善未来的观察。模型在闭环控制下直接从低分辨率RGB图像预测关节指令。结果表明,低分辨率全景视觉足以可靠完成任务,且相对关节增量预测显著优于绝对位置预测。这些结果证明,在可复现的设置中,视觉引导的主动感知可通过行为克隆自然涌现。
原文摘要 · Abstract (English)
We investigate whether behavior cloning is sufficient to produce active perception in a structured object-finding task. A low-cost robot arm equipped with a wrist-mounted egocentric RGB camera must reposition to center a partially visible plant before triggering a grasp signal, requiring actions that improve future observations. The model predicts joint commands directly from low-resolution RGB images under closed-loop control. We show that low-resolution egocentric vision is sufficient for reliable task completion and that predicting relative joint deltas substantially outperforms absolute joint position prediction in our setting. These results demonstrate that visually grounded active perception can emerge from behavior cloning in a reproducible setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。