构建大规模3D多人多物交互数据集,支持复杂互动建模与动作识别。
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
- 提出双补丁结构表示法,联合建模物体与交互关系。
- 在MMHOI和CORE4D上实现最先进的3D姿态与动作预测精度。
- 适合研究复杂人机交互、三维场景理解的科研人员使用。
现实场景中常存在多人与多物之间的因果性、目标导向或协作式交互,但现有3D人-物交互(HOI)基准仅涵盖其中一小部分。为填补这一空白,我们提出MMHOI——一个大规模的多人群体多物交互数据集,包含12种日常场景图像。该数据集对每个人和每个物体提供完整的3D形状与姿态标注,并标注了78类动作和14个交互特异性身体部位,构成下一代HOI研究的全面测试平台。基于此,我们提出MMHOI-Net,一种端到端的基于Transformer的神经网络,用于联合估计人-物3D几何结构、交互关系及对应动作。框架的核心创新在于采用结构化的双补丁表示来建模物体及其交互,并结合动作识别增强交互预测能力。在MMHOI和近期提出的CORE4D数据集上的实验表明,该方法在多主体人-物交互建模中达到当前最优性能,兼具高准确率与优良重建质量。MMHOI数据集已公开于https://zenodo.org/records/17711786。
原文摘要 · Abstract (English)
Real-world scenes often feature multiple humans interacting with multiple objects in ways that are causal, goal-oriented, or cooperative. Yet existing 3D human-object interaction (HOI) benchmarks consider only a fraction of these complex interactions. To close this gap, we present MMHOI -- a large-scale, Multi-human Multi-object Interaction dataset consisting of images from 12 everyday scenarios. MMHOI offers complete 3D shape and pose annotations for every person and object, along with labels for 78 action categories and 14 interaction-specific body parts, providing a comprehensive testbed for next-generation HOI research. Building on MMHOI, we present MMHOI-Net, an end-to-end transformer-based neural network for jointly estimating human-object 3D geometries, their interactions, and associated actions. A key innovation in our framework is a structured dual-patch representation for modeling objects and their interactions, combined with action recognition to enhance the interaction prediction. Experiments on MMHOI and the recently proposed CORE4D datasets demonstrate that our approach achieves state-of-the-art performance in multi-HOI modeling, excelling in both accuracy and reconstruction quality. The MMHOI dataset is publicly available at https://zenodo.org/records/17711786.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。