arXiv:2510.06339cs.RO2025-10被引 4

视觉引导粗略抓取,触觉反馈精细调控,实现无需模型的物体操作。

Vi-TacMan: Articulated Object Manipulation via Vision and Touch

  • 用视觉提供抓取位置和大致方向,作为触觉控制的起点。
  • 在超过5万次仿真与真实测试中,成功率显著高于基线(全部p<0.0001)。
  • 无需预先知道物体结构,适合复杂多变的真实环境应用。

机器人在人类环境中自主操作活动物体仍是核心挑战。基于视觉的方法可推断隐藏运动学,但对陌生物体估计不准;触觉方法通过接触反馈实现鲁棒控制,但需精确初始化。这提示了天然互补性:视觉提供全局引导,触觉实现局部精确。然而,现有框架未系统利用这一协同机制实现通用活动物体操作。本文提出Vi-TacMan,利用视觉提出抓取位置和粗略方向,作为触觉控制器的初始输入。通过引入表面法向量作为几何先验,并采用冯·米塞斯-费舍尔分布建模方向,方法在多个基准上显著优于基线(所有p<0.0001)。关键在于,操作无需显式运动学模型——触觉控制器通过实时接触调节,不断优化视觉提供的粗略估计。在超过5万次模拟及多样化真实物体上的测试验证了跨类别泛化能力。本工作表明,仅需粗略视觉线索配合触觉反馈,即可实现可靠操作,为非结构化环境中自主系统提供了可扩展范式。

原文摘要 · Abstract (English)

Autonomous manipulation of articulated objects remains a fundamental challenge for robots in human environments. Vision-based methods can infer hidden kinematics but can yield imprecise estimates on unfamiliar objects. Tactile approaches achieve robust control through contact feedback but require accurate initialization. This suggests a natural synergy: vision for global guidance, touch for local precision. Yet no framework systematically exploits this complementarity for generalized articulated manipulation. Here we present Vi-TacMan, which uses vision to propose grasps and coarse directions that seed a tactile controller for precise execution. By incorporating surface normals as geometric priors and modeling directions via von Mises-Fisher distributions, our approach achieves significant gains over baselines (all p<0.0001). Critically, manipulation succeeds without explicit kinematic models -- the tactile controller refines coarse visual estimates through real-time contact regulation. Tests on more than 50,000 simulated and diverse real-world objects confirm robust cross-category generalization. This work establishes that coarse visual cues suffice for reliable manipulation when coupled with tactile feedback, offering a scalable paradigm for autonomous systems in unstructured environments.

机器人操作视觉触觉融合无模型控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。