通过动态调整视角和焦距,提升机器人操作的精准感知能力。
AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization
- 结合头戴设备与可调云台,实时优化观察视角和放大倍数。
- 仿真中任务成功率提升5%-17%,真实场景提升超25%。
- 适合需要高精度视觉引导的复杂环境机器人操作任务。
在复杂场景中进行机器人操作需要对任务相关细节进行精确感知,但固定或次优的视角常导致细粒度感知受限并引发遮挡,制约了模仿学习策略的表现。我们提出AVR(主动视觉驱动机器人系统),一个双臂遥操作与学习框架,将头戴设备追踪的视角控制(HMD到2自由度云台)与电动变焦镜头结合,在数据采集和部署过程中始终使目标保持居中且处于合适尺度。在仿真中,通过AVR插件增强RoboTwin演示,模拟主动视觉(基于感兴趣区域的视角变化、保持纵横比的裁剪及显式变焦比例、超分辨率),在多种操作任务中实现5%-17%的任务成功率提升。在真实平台测试中,大多数任务的成功率显著提高,相比静态视角基线提升超过25%;进一步研究表明,该方法在遮挡、杂乱环境和光照变化下仍具鲁棒性,并能泛化至未见过的环境与物体。这些成果为追求人类级灵巧性与精度的机器人精密操作方法铺平了道路。
原文摘要 · Abstract (English)
Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present AVR (Active Vision-driven Robotics), a bimanual teleoperation and learning framework that unifies head-tracked viewpoint control (HMD-to-2-DoF gimbal) with motorized optical zoom to keep targets centered at an appropriate scale during data collection and deployment. In simulation, an AVR plugin augments RoboTwin demonstrations by emulating active vision (ROI-conditioned viewpoint change, aspect-ratio-preserving crops with explicit zoom ratios, and super-resolution), yielding 5-17% gains in task success across diverse manipulations. On our real-world platform, AVR improves success on most tasks, with over 25% gains compared to the static-view baseline, and extended studies further demonstrate robustness under occlusion, clutter, and lighting disturbances, as well as generalization to unseen environments and objects. These results pave the way for future robotic precision manipulation methods in the pursuit of human-level dexterity and precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。