arXiv:2604.21331cs.RO2026-04

用指尖摄像头实现多视角视觉感知,提升机器人灵巧操作能力。

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception

论文配图:FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception
图 1 · 摘自论文原文
  • 在每根手指装微型摄像头,获取手部与环境的多视角图像。
  • 基于扩散模型从人类示范中学习复杂操作技能,成功率达80.8%。
  • 适合需要精准触觉与视觉协同的机器人灵巧操作研究者。

当前灵巧操作通常依赖单一腕部视角,常被遮挡且限制性能。本文提出FingerViP,一种利用指尖视觉感知的视觉-运动策略系统。我们设计了集成微型相机的视觉增强指尖模块,并安装于多指机械手上。指尖摄像头提供了手部及周围环境的全面多视角反馈,显著提升了视觉感知能力。在此基础上,我们构建了一个以第三视角相机和多视角指尖视觉为条件的扩散模型全身体感运动策略,直接从人类示范中学习复杂操作技能。为增强视角-本体感觉对齐与接触感知,每个指尖视觉特征均融合了对应相机位姿编码和各指关节电流编码。我们在多个真实世界挑战任务中验证了多视角指尖视觉的有效性,包括在密闭盒内按按钮、从不稳支撑中取棒、绕过遮挡帘取物,以及长时程柜门开启与物体检索,整体成功率高达80.8%。所有硬件设计与代码将完全开源。

原文摘要 · Abstract (English)

The current practice of dexterous manipulation generally relies on a single wrist-mounted view, which is often occluded and limits performance on tasks requiring multi-view perception. In this work, we present FingerViP, a learning system that utilizes a visuomotor policy with fingertip visual perception for dexterous manipulation. Specifically, we design a vision-enhanced fingertip module with an embedded miniature camera and install the modules on each finger of a multi-fingered hand. The fingertip cameras substantially improve visual perception by providing comprehensive, multi-view feedback of both the hand and its surrounding environment. Building on the integrated fingertip modules, we develop a diffusion-based whole-body visuomotor policy conditioned on a third-view camera and multi-view fingertip vision, which effectively learns complex manipulation skills directly from human demonstrations. To improve view-proprioception alignment and contact awareness, each fingertip visual feature is augmented with its corresponding camera pose encoding and per-finger joint-current encoding. We validate the effectiveness of the multi-view fingertip vision and demonstrate the robustness and adaptability of FingerViP on various challenging real-world tasks, including pressing buttons inside a confined box, retrieving sticks from an unstable support, retrieving objects behind an occluding curtain, and performing long-horizon cabinet opening and object retrieval, achieving an overall success rate of 80.8%. All hardware designs and code will be fully open-sourced.

灵巧操作视觉感知机器人扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。