arXiv:2601.20835cs.CVcs.AI2026-01被引 2

让3D人物按功能与场景互动,支持任意任务提示。

Open-Vocabulary Functional 3D Human-Scene Interaction Generation

  • 基于功能语义推理识别场景中可交互元素并构建接触图
  • 通过视觉语言模型生成人体姿态并分阶段优化确保合理
  • 无需训练即可实现细粒度功能交互,适合内容创作与机器人模拟

生成能功能性地与3D场景互动的3D人物仍是开放难题,应用于具身AI、机器人及交互内容创作。核心挑战在于理解场景中功能元素的语义及其对应的3D人体姿态。现有方法通常缺乏对物体功能和人-场景接触的显式推理,导致交互不自然或功能错误。本文提出FunHSI,一种无需训练的功能驱动框架,可基于开放词汇任务提示生成功能正确的交互。给定任务提示后,该框架执行功能感知的接触推理,重建场景几何,通过接触图建模高层交互;随后利用视觉语言模型合成图像中执行任务的人体,并估计3D躯干与手部姿态;最后通过分阶段优化细化3D人体配置,确保物理合理性与功能正确性。相比现有方法,FunHSI不仅能生成更合理的通用交互(如“坐在沙发上”),还能支持细粒度功能交互(如“提高房间温度”)。大量实验表明,FunHSI在多样化的室内外场景中均能稳定生成功能正确且物理合理的3D人-场景交互。

原文摘要 · Abstract (English)

Generating 3D humans that functionally interact with 3D scenes remains an open problem with applications in embodied AI, robotics, and interactive content creation. The key challenge involves reasoning about both the semantics of functional elements in 3D scenes and the 3D human poses required to achieve functionality-aware interaction. Unfortunately, existing methods typically lack explicit reasoning over object functionality and the corresponding human-scene contact, resulting in implausible or functionally incorrect interactions. In this work, we propose FunHSI, a training-free, functionality-driven framework that enables functionally correct human-scene interactions from open-vocabulary task prompts. Given a task prompt, FunHSI performs functionality-aware contact reasoning to identify functional scene elements, reconstruct their 3D geometry, and model high-level interactions via a contact graph. We then leverage vision-language models to synthesize a human performing the task in the image and estimate proposed 3D body and hand poses. Finally, the proposed 3D body configuration is refined via stage-wise optimization to ensure physical plausibility and functional correctness. In contrast to existing methods, FunHSI not only synthesizes more plausible general 3D interactions, such as "sitting on a sofa'', while supporting fine-grained functional human-scene interactions, e.g., "increasing the room temperature''. Extensive experiments demonstrate that FunHSI consistently generates functionally correct and physically plausible human-scene interactions across diverse indoor and outdoor scenes.

3D生成人机交互功能推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。