用文字指令生成符合功能需求的3D手物交互,无需标注数据
FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidance
- 分两阶段生成:先根据文字生成手物3D模型,再优化姿态
- 在无额外标注条件下实现精准且物理合理的功能抓握
- 适合机器人抓取、虚拟现实等需要语义化交互的场景
手物交互(HOI)是人与环境交互的基础,但其精细复杂的姿态对手势控制构成重大挑战。尽管人工智能与机器人技术取得显著进展,使机器能够理解并模拟手物交互,但捕捉功能性抓握的语义仍面临巨大困难。以往方法虽能生成稳定正确的3D抓握,却因忽略抓握语义而难以实现功能性抓握。为此,我们提出创新的两阶段框架——功能抓握合成网络(FGS-Net),通过功能文本驱动生成3D HOI。该框架包含文本引导的3D模型生成器(FGG)和姿态优化策略(FGR)。FGG根据文本输入生成手与物体的3D模型,FGR利用物体姿态近似器和能量函数精调姿态,确保手物相对位置符合人类意图且物理合理。大量实验表明,本方法在无需额外3D标注数据的情况下,实现了精确且高质量的HOI生成。
原文摘要 · Abstract (English)
Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenges for gesture control. Despite significant advances in AI and robotics, enabling machines to understand and simulate hand-object interactions, capturing the semantics of functional grasping tasks remains a considerable challenge. While previous work can generate stable and correct 3D grasps, they are still far from achieving functional grasps due to unconsidered grasp semantics. To address this challenge, we propose an innovative two-stage framework, Functional Grasp Synthesis Net (FGS-Net), for generating 3D HOI driven by functional text. This framework consists of a text-guided 3D model generator, Functional Grasp Generator (FGG), and a pose optimization strategy, Functional Grasp Refiner (FGR). FGG generates 3D models of hands and objects based on text input, while FGR fine-tunes the poses using Object Pose Approximator and energy functions to ensure the relative position between the hand and object aligns with human intent and remains physically plausible. Extensive experiments demonstrate that our approach achieves precise and high-quality HOI generation without requiring additional 3D annotation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。