arXiv:2608.01905cs.CV2026-08

从一张照片和文字指令生成3D手物交互动画,无需预设物体模型。

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

论文配图:PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph
图 1 · 摘自论文原文
  • 用视觉语言模型解析图像与指令,生成结构化任务描述。
  • 在真实照片上实现高精度接触、低穿透,支持未见物体和开放词汇。
  • 通过学习可迁移的抓握先验,让手部动作自然且符合物理约束。

手物交互(HOI)是增强现实、数字人和具身交互中的基础行为。现有方法通常依赖预定义物体几何或特定任务条件,难以处理真实世界输入。为此,本文提出从单张RGB照片和开放词汇语言指令中合成3D手物交互序列的PhotoHOI。首先,利用视觉语言模型将输入图像与指令解析为包含交互对象、目标区域和空间关系的结构化任务规范;随后,基于恢复的物体状态、支撑关系和场景几何,规划平滑且避碰的物体运动轨迹;为实现对真实照片和未知物体的泛化,模型从大规模可用性与HOI数据中学习可迁移的任务条件接触与接触条件抓握先验,并在学习的隐空间中精炼抓握姿态,约束优化于合理手姿流形内。在GRAB与H2O数据集上的实验表明,相比代表性基线,接触质量更高、穿透更少;在真实照片上的结果进一步验证了更高的任务成功率与场景一致性,支持未见物体与开放词汇指令的泛化能力。

原文摘要 · Abstract (English)

Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.

3D生成手物交互视觉语言模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。