让虚拟人物在真实场景中精准操作物体并自由移动
HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene Perception
- 分层感知场景,先规划路径再生成动作
- 生成长序列动作且无穿透,手部操作精确到指节
- 适合影视动画与虚拟人开发,无需人工调整
生成高保真全身体态与动态物体及静态场景的交互仍是计算机图形学与动画领域的关键挑战。现有方法常忽略场景上下文,导致不合理的穿透;而人体-场景交互方法难以协调精细操作与远距离导航。为此,我们提出HOSIG框架,通过分层场景感知合成全身体交互。方法分解为三部分:1)场景感知抓取姿态生成器,结合局部几何约束实现无碰撞全身姿态与精准手物接触;2)启发式导航算法,利用压缩2D地板图与双组件空间推理,在复杂室内环境自主规划避障路径;3)场景引导运动扩散模型,通过空间锚点与双空间无分类器引导,生成轨迹控制、指节级精度的全身动作。TRUMANS数据集上的大量实验表明,该框架优于当前最优方法。其支持自回归生成无限长度动作,且几乎无需人工干预。本工作弥合了场景感知导航与灵巧物体操作之间的关键鸿沟,推动具身交互合成前沿发展。代码将于发表后公开。项目页面:http://yw0208.github.io/hosig
原文摘要 · Abstract (English)
Generating high-fidelity full-body human interactions with dynamic objects and static scenes remains a critical challenge in computer graphics and animation. Existing methods for human-object interaction often neglect scene context, leading to implausible penetrations, while human-scene interaction approaches struggle to coordinate fine-grained manipulations with long-range navigation. To address these limitations, we propose HOSIG, a novel framework for synthesizing full-body interactions through hierarchical scene perception. Our method decouples the task into three key components: 1) a scene-aware grasp pose generator that ensures collision-free whole-body postures with precise hand-object contact by integrating local geometry constraints, 2) a heuristic navigation algorithm that autonomously plans obstacle-avoiding paths in complex indoor environments via compressed 2D floor maps and dual-component spatial reasoning, and 3) a scene-guided motion diffusion model that generates trajectory-controlled, full-body motions with finger-level accuracy by incorporating spatial anchors and dual-space classifier-free guidance. Extensive experiments on the TRUMANS dataset demonstrate superior performance over state-of-the-art methods. Notably, our framework supports unlimited motion length through autoregressive generation and requires minimal manual intervention. This work bridges the critical gap between scene-aware navigation and dexterous object manipulation, advancing the frontier of embodied interaction synthesis. Codes will be available after publication. Project page: http://yw0208.github.io/hosig
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。