arXiv:2607.28394cs.CV2026-07

系统梳理大模型如何提升手物交互建模的准确性与泛化能力

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

论文配图:Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
图 1 · 摘自论文原文
  • 构建六类任务与八类先验知识的分类体系
  • 揭示几何、语义、视觉三类先验在流程中的注入方式
  • 适合机器人学习与交互建模研究者参考

手物交互(HOI)建模因需联合推理手部关节、物体几何、接触关系、语义信息及动态特性,且面临严重视觉不确定性而极具挑战。基础模型通过大规模跨域数据学习可迁移先验知识,为突破任务特定数据与模型的局限提供了新路径。然而现有研究文献分散,多仅笼统描述为“使用大模型”,缺乏对引入何种知识、知识在HOI流程中何处切入、缓解何种不确定性等关键问题的系统分析。本文首次对用于HOI的基础模型先验进行系统综述,将文献分为六类任务(涵盖重建与生成)。更重要的是,提出包含八种子先验的分类体系:几何类(形状检索、重建、空间重建)、语义类(语义定位、语言推理)、视觉类(视觉表示、图像生成、视频生成)。基于此分类,系统分析不同先验在流程中的表征、注入与适配方式。除探讨基础模型如何赋能HOI外,进一步考察HOI知识在机器人学习中的应用,包括人类数据预训练、人到机器人的技能迁移、以及基于HOI的机器人数据生成。最后总结常用数据集与评估协议,讨论当前局限并展望未来方向。为支持长期进展,我们维护一个持续更新的方法与基准聚合仓库。

原文摘要 · Abstract (English)

Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.

手物交互大模型机器人学习先验知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。