修复多人协作抓取动作中的手物错位与抖动问题,提升动作自然性。
MOCHI: Motion Enhancement of Collaborative Human-object Interactions

- 先优化手部抓握姿态,再用扩散模型优化全身动作
- 解决手物对齐、运动抖动和手指细节缺失等噪声问题
- 适合做多人体交互数据增强或动画关键帧生成
多人协作抓取动作具有动态复杂的特点,需参与者与共享物体间持续互预与调整。高质量数据采集是建模的基础,但因人-人与人-物交互同时发生,难以实现。现有数据常含多种伪影:手物接触错位、序列抖动与时间不一致、指节级关节信息缺失。为此,我们提出MOCHI(Motion Enhancement of Collaborative Human-object Interactions),一个两阶段的噪声数据增强框架。首先基于噪声身体输入,通过优化生成物理合理且语义一致的手部抓握姿态,并扩展为完整的手物交互序列;随后,利用单人运动先验的扩散去噪框架,对所有参与者全身体态进行精炼。优化过程中,将人-物与人-人交互信息编码进单人先验中。实验表明,该方法在多种真实采集及生成的MHOI数据上均有效,具备跨参与者数量与交互类型鲁棒性,支持基于关键帧的交互生成与物体几何变化下的数据增强。
原文摘要 · Abstract (English)
Collaborative human-object interaction shows dynamic and complex movements that require mutual anticipation and continuous adjustment between participants and the shared object. Modeling such collaborative multi-human object interaction (MHOI) scenarios requires high-quality data acquisition as a foundational step; however, this is challenging due to the inherent complexity of MHOI where human-human and human-object interactions occur simultaneously. Such complexity leads to noisy MHOI captures characterized by several artifacts: contact misalignment between hands and objects, motion jitter and temporal inconsistencies in the captured sequences, and missing or incomplete finger-level articulation details. To address these challenges, we present MOCHI (MOtion Enhancement of Collaborative Human-object Interactions), a two-stage framework for enhancing noisy MHOI data. Our approach first generates physically plausible hand grasps through optimization from noisy body input, producing grasps that are both physically plausible and semantically consistent with the body pose, where these optimized grasps are extended into complete hand-object interaction sequences. Consequently, the full-body motion for all participants are refined through a diffusion-based noise optimization framework that uses single-person motion priors. During the optimization process, we introduce optimization objectives to encode human-object and human-human interaction information within these single-person priors. Experimental results demonstrate the effectiveness of our pipeline across diverse MHOI data, either acquired by existing capture methods or synthesized by generative models. We further show robustness of our system across varying numbers of participants and types of interactions, and demonstrate various applications including keyframe-based MHOI creation and data augmentation through varying object geometries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。