arXiv:2604.04843cs.CVcs.AI2026-04

通过动态感知与迭代优化,实现真实感人-物-场景交互生成。

InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement

  • 采用分步细化框架,结合一致性模型的去噪过程逐步生成交互。
  • 在未标注数据少的情况下仍达顶尖性能,支持新场景泛化。
  • 无需精细几何信息即可避免碰撞,适合实时应用与虚拟仿真。

人-物-场景交互(HOSI)生成在具身智能、模拟与动画中有广泛应用。与人-物交互(HOI)和人-场景交互(HSI)不同,HOSI需处理动态物体与场景变化,但标注数据稀缺。为此,我们提出一种从粗到精的指令条件化交互生成框架,明确对齐一致性模型的迭代去噪过程。特别地,引入动态感知策略,利用前一轮优化的轨迹更新场景上下文,并在每一步去噪中调节后续优化,实现一致交互。为减少物理错误,设计了抗碰撞引导机制,在不依赖精细场景几何的前提下有效缓解碰撞与穿透,支持实时生成。针对数据稀缺问题,提出混合训练策略:通过向HOI数据集注入体素化场景占据信息合成伪HOSI样本,并与高质量HSI数据联合训练,实现交互学习的同时保持真实场景感知。大量实验表明,该方法在HOSI与HOI生成上均达到当前最优性能,并具备强泛化能力至未见场景。

原文摘要 · Abstract (English)

Human-object-scene interactions (HOSI) generation has broad applications in embodied AI, simulation, and animation. Unlike human-object interaction (HOI) and human-scene interaction (HSI), HOSI generation requires reasoning over dynamic object-scene changes, yet suffers from limited annotated data. To address these issues, we propose a coarse-to-fine instruction-conditioned interaction generation framework that is explicitly aligned with the iterative denoising process of a consistency model. In particular, we adopt a dynamic perception strategy that leverages trajectories from the preceding refinement to update scene context and condition subsequent refinement at each denoising step of consistency model, yielding consistent interactions. To further reduce physical artifacts, we introduce a bump-aware guidance that mitigates collisions and penetrations during sampling without requiring fine-grained scene geometry, enabling real-time generation. To overcome data scarcity, we design a hybrid training startegy that synthesizes pseudo-HOSI samples by injecting voxelized scene occupancy into HOI datasets and jointly trains with high-fidelity HSI data, allowing interaction learning while preserving realistic scene awareness. Extensive experiments demonstrate that our method achieves state-of-the-art performance in both HOSI and HOI generation, and strong generalization to unseen scenes. Project page: https://yudezou.github.io/InfBaGel-page/

交互生成一致性模型动态感知虚拟仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。