高精度4D人体物体重叠数据集,助力动画与机器人交互建模
HUMOTO: A 4D Dataset of Mocap Human Object Interactions
- 用场景驱动的LLM脚本生成有目的性的自然动作序列
- 包含735段、7875秒的真人动捕数据,覆盖63种物体和72个关节
- 适合做动作生成、机器人交互与具身智能研究的团队使用
我们提出人类与物体互动的高保真数据集HUMOTO,用于运动生成、计算机视觉与机器人应用。该数据集包含735个序列(共7,875秒,30帧/秒),涵盖63个精确建模的物体和72个刚性部件。创新之处在于采用场景驱动的LLM脚本生成流程,实现具有自然进展的完整任务;并结合动捕与多摄像头系统有效处理遮挡问题。数据覆盖从烹饪到户外野餐等多样活动,同时保证物理准确性和任务逻辑连贯性。专业艺术家对每段数据进行严格清洗与验证,显著减少脚部滑移和物体穿透。我们还提供了与其他数据集的基准对比。HUMOTO通过全身体运动与多物体同步交互,解决关键数据采集难题,为动画、机器人及具身智能系统提供实用支持。
原文摘要 · Abstract (English)
We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated parts. Our innovations include a scene-driven LLM scripting pipeline creating complete, purposeful tasks with natural progression, and a mocap-and-camera recording setup to effectively handle occlusions. Spanning diverse activities from cooking to outdoor picnics, HUMOTO preserves both physical accuracy and logical task flow. Professional artists rigorously clean and verify each sequence, minimizing foot sliding and object penetrations. We also provide benchmarks compared to other datasets. HUMOTO's comprehensive full-body motion and simultaneous multi-object interactions address key data-capturing challenges and provide opportunities to advance realistic human-object interaction modeling across research domains with practical applications in animation, robotics, and embodied AI systems. Project: https://jiaxin-lu.github.io/humoto/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。