arXiv:2503.16289cs.CV2025-03ICCV被引 16

用场景感知的插帧技术,让人体与环境互动更自然真实。

SceneMI: Motion In-betweening for Modeling Human-Scene Interactions

  • 通过双场景描述符捕捉全局与局部环境信息
  • 在噪声数据上仍能生成高质量动作,适用于真实传感器数据
  • 适合动画师控制角色动作或修复低质量动作数据

建模人体-场景交互(HSI)对理解日常行为至关重要。现有生成方法虽有进展,但在实际应用中可控性与灵活性不足。为此,我们提出将HSI建模重构为场景感知的动作插帧问题,更具可操作性。本文提出SceneMI框架,支持多种实用场景,包括3D场景中关键帧引导的角色动画和提升不完整HSI数据的运动质量。SceneMI采用双场景描述符全面编码全局与局部场景上下文,并利用扩散模型固有的去噪特性,实现对噪声关键帧的泛化能力。实验表明,SceneMI在场景感知插帧任务中表现优异,且在真实世界数据集GIMO上成功应用,该数据集由带噪声的惯性传感器和智能手机采集。此外,我们还展示了其从单目视频重建HSI的能力。

原文摘要 · Abstract (English)

Modeling human-scene interactions (HSI) is essential for understanding and simulating everyday human behaviors. Recent approaches utilizing generative modeling have made progress in this domain; however, they are limited in controllability and flexibility for real-world applications. To address these challenges, we propose reformulating the HSI modeling problem as Scene-aware Motion In-betweening - a more tractable and practical task. We introduce SceneMI, a framework that supports several practical applications, including keyframe-guided character animation in 3D scenes and enhancing the motion quality of imperfect HSI data. SceneMI employs dual scene descriptors to comprehensively encode global and local scene context. Furthermore, our framework leverages the inherent denoising nature of diffusion models to generalize on noisy keyframes. Experimental results demonstrate SceneMI's effectiveness in scene-aware keyframe in-betweening and generalization to the real-world GIMO dataset, where motions and scenes are acquired by noisy IMU sensors and smartphones. We further showcase SceneMI's applicability in HSI reconstruction from monocular videos.

动作生成场景感知扩散模型数据修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。