arXiv:2601.08246cs.RO2026-01被引 2

用人类视频和扩散模型生成精细抓握,无需真实机器人数据

FSAG: Enhancing Human-to-Dexterous-Hand Finger-Specific Affordance Grounding via Diffusion Models

  • 从人类抓握视频提取语义接触区域,结合深度图构建抓握目标
  • 在多种物体和手型上实现稳定、自然的多指抓握,泛化性强
  • 仅需单目深度图+预训练模型,适合无硬件依赖的智能抓取

灵巧抓握合成需同时满足功能意图与物理可行性,但现有方法常将语义理解与优化分离,导致在物体和姿态变化下产生不稳定或非功能性接触。这一挑战因多指手的高维性与运动多样性而加剧,许多方法依赖昂贵的仿真或真实世界采集的大规模硬件专用抓握数据集。本文提出一种数据高效框架,通过利用预训练生成式扩散模型中的对象中心语义先验,绕过机器人抓握数据收集。从原始人类视频中提取时间对齐、细粒度的抓握语义先验,并与深度图像中的3D场景几何融合,推断出语义驱动的接触目标。进一步将这些先验区域引入抓握优化目标,在优化过程中显式引导每个指尖向预测区域靠近。所提系统在常见物体与工具上生成稳定、类人化的多接触抓握,且对类别内未见过的物体实例、姿态变化及多种手型具有强泛化能力。本工作(i)引入基于视觉-语言生成先验的语义先验提取流程用于灵巧抓握;(ii)证明无需构建硬件专用抓握数据集即可实现跨手型泛化;(iii)表明单个深度模态结合基础模型语义即可实现高性能抓握合成。结果揭示了一条由人类示范和预训练生成模型驱动的可扩展、硬件无关灵巧操作路径。

原文摘要 · Abstract (English)

Dexterous grasp synthesis must jointly satisfy functional intent and physical feasibility, yet existing pipelines often decouple semantic grounding from refinement, yielding unstable or non-functional contacts under object and pose variations. This challenge is exacerbated by the high dimensionality and kinematic diversity of multi-fingered hands, which makes many methods rely on large, hardware-specific grasp datasets collected in simulation or through costly real-world trials. We propose a data-efficient framework that bypasses robot grasp data collection by exploiting object-centric semantic priors in pretrained generative diffusion models. Temporally aligned and fine-grained grasp affordances are extracted from raw human video demonstrations and fused with 3D scene geometry from depth images to infer semantically grounded contact targets. We further incorporate these affordance regions into the grasp refinement objective, explicitly guiding each fingertip toward its predicted region during optimization. The resulting system produces stable, human-intuitive multi-contact grasps across common objects and tools, while exhibiting strong generalization to previously unseen object instances within a category, pose variations, and multiple hand embodiments.This work (i) introduces a semantic affordance extraction pipeline leveraging vision--language generative priors for dexterous grasping, (ii) demonstrates cross-hand generalization without constructing hardware-specific grasp datasets, and (iii) establishes that a single depth modality suffices for high-performance grasp synthesis when coupled with foundation-model semantics. Our results highlight a path toward scalable, hardware-agnostic dexterous manipulation driven by human demonstrations and pretrained generative models.

灵巧抓握扩散模型语义先验零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。