构建首个大规模多模态人机交互分析基准,支持生成与感知双重任务。
Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

- 基于新型混合动捕系统采集1.1万段高保真人际交互数据。
- 涵盖810万帧数据,含手指精细动作与多层次语义标注。
- 提出统一建模框架OpenHHI,实现交互理解与生成的协同优化。
感知与合成人际交互是构建智能数字人系统的基础。然而现有数据集与建模方法受限于低精度运动捕捉、缺乏灵巧手势及丰富的多模态标注,且交互表示碎片化、评估协议不统一,阻碍了公平严谨的基准测试。为此,我们提出Inter-X++,一个全面且大规模的基准,旨在赋能多样化的交互分析。通过新型混合动作捕捉系统,Inter-X++包含11,388个高保真交互序列和超过810万帧数据,精确呈现全身运动与手指关节细节。同时,我们丰富数据标注维度,涵盖分层细粒度文本描述、交互类别、因果顺序、人物关系与个性属性,以及顶点级接触图与物理正则化约束。基于这些详尽标注,我们构建了覆盖四类下游任务的统一测试平台,对齐生成与感知范式。为消除评测歧义,我们系统性标准化交互表示与评估协议。进一步地,我们提出OpenHHI——一个统一的人际交互表示与建模范式,联合优化交互重建与语义理解。大量实验表明,OpenHHI在生成与感知任务上均达到顶尖性能,证实该统一表示能有效融合交互理解与生成。
原文摘要 · Abstract (English)
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent evaluation protocols also impede fair and rigorous benchmarking. To systematically address these bottlenecks, we present Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis. Captured via a novel hybrid motion capture system, Inter-X++ provides 11,388 high-fidelity interaction sequences and over 8.1M frames, featuring precise whole-body movements and detailed finger articulations. Meanwhile, we enrich the data foundation with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction orders, the relationship and personality of the subjects, as well as vertex-level contact maps and physically regularized constraints. Leveraging these elaborate annotations, we formulate a unified testing ground comprising four categories of downstream tasks that symmetrically span both generative and perceptive paradigms. To eliminate benchmarking ambiguities, we systematically standardize the interaction representations and evaluation protocols. Finally, we go beyond dataset construction to propose OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments reveal that OpenHHI achieves state-of-the-art performance on both generation and perception tasks. This definitively proves that our unified representation successfully bridges interaction understanding and generation simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。