首个宠物与人交互的多模态4D数据集,用于生成真实互动动作。
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

- 构建同步多视角系统采集人狗互动视频与标注数据。
- 含680万帧、13只不同犬种的互动数据,支持3D姿态与网格重建。
- 提出新框架生成逼真动作,显著优于基线模型。
由于缺乏高质量大规模数据集,人宠交互估计与生成仍处于探索阶段。我们提出了InterPet4D,首个捕捉自然人狗互动的多模态4D数据集。通过同步多视角采集系统,记录了人狗服从任务,并提供双主体标注:包括多视角与第一人称视频、分割图、2D/3D关键点、网格模型及音频。数据集包含13只来自11个犬种的狗与23名人类参与者的680万帧数据。我们进一步提出InterPetMoGen框架用于人宠交互动作生成。所提模型在生成质量上达到FID 11.21,显著优于Seq2Seq和DiT基线,证明了InterPet4D在建模真实人宠互动中的有效性。
原文摘要 · Abstract (English)
Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between humans and dogs. Using a synchronized multi-view capture system, we record human-dog obedience tasks and provide annotations for both humans and dogs, including multi-view and egocentric videos, segmentations, 2D and 3D keypoints, meshes, and audio tracks. InterPet4D consists of 6.8 million frames collected from 13 dogs of 11 breeds interacting with 23 human participants. We further introduce the InterPetMoGen framework for human-pet interaction motion generation. Our proposed model achieves an FID score of 11.21 and substantially outperforms the Seq2Seq and DiT baselines, demonstrating the effectiveness of InterPet4D for modeling realistic human-pet interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。