用动态超图建模多主体交互,提升精细操作识别准确率
Beyond Pairwise Relations: Dynamic Manipulation Hypergraphs for Vision-Based Human Activity Recognition

- 用超边表示手、物、工具等多主体的高阶关系,而非简单配对
- 在EPIC-KITCHENS-100/VISOR上比配对图提升6.9个百分点
- 能自动识别关键操作阶段,适合动作理解与机器人学习
细粒度操作识别需要建模手、物体、工具和支撑面之间随时间演化的复杂关系。传统图方法仅使用二元边,将协调事件拆解为孤立的成对关系。本文提出动态操纵超图框架,将多实体配置表示为高阶关系单元。每个时间步,通过外观、空间、运动和语义角色特征编码相关实体。利用邻近性、接触和运动耦合谓词生成并排序超边候选。超图推理网络执行节点到超边、超边到节点的消息传递,并在演化交互结构上应用时间注意力。该框架提供无类别超边重要性评分,可识别模型关注的实体配置和时间区间,但不作为因果解释。在EPIC-KITCHENS-100/VISOR和Assembly101数据集上,基于标注辅助实体定位协议进行定量评估。视频仅输入和基于实体的方法提供上下文对比,匹配的配对图和静态超图作为主要对照基线,因使用相同实体输入和可比关系设置。所提方法在EPIC-KITCHENS-100/VISOR上相比匹配配对图提升HO-F1 6.9个百分点,在Assembly101上提升9.5个百分点;相比静态超图分别提升4.4和5.8个百分点。在ARCTIC上的定性分析显示,高排名超边与接触密集的操作阶段高度对应。结果表明,时变高阶关系建模对细粒度操作识别具有显著价值。
原文摘要 · Abstract (English)
Fine-grained manipulation recognition requires modeling evolving relations among hands, objects, tools, and supporting surfaces. Conventional graph-based methods use pairwise edges that can fragment a coordinated event into disconnected binary relations. We propose a dynamic manipulation hypergraph framework that represents multi-entity configurations as higher-order relational units. At each temporal step, relevant entities are encoded using appearance, spatial, motion, and semantic-role features. Hyperedge candidates are instantiated and ranked using proximity, contact, and motion-coupling predicates. A hypergraph reasoning network performs node-to-hyperedge and hyperedge-to-node message passing, followed by temporal attention over the evolving interaction structure. The framework provides class-agnostic hyperedge-importance scores that identify entity configurations and temporal intervals emphasized by the model without treating them as causal explanations. Quantitative evaluation is conducted on EPIC-KITCHENS-100/VISOR and Assembly101 under an annotation-assisted entity-localization protocol. Video-only and entity-based methods provide contextual comparisons, while a matched pairwise graph and a static hypergraph serve as the principal controlled baselines because they use identical entity inputs and comparable relational settings. The proposed method improves HO-F1 over the matched pairwise graph by 6.9 percentage points on EPIC-KITCHENS-100/VISOR and 9.5 points on Assembly101, and exceeds the static hypergraph by 4.4 and 5.8 points, respectively. Qualitative analysis on ARCTIC further shows correspondence between highly ranked hyperedges and contact-rich manipulation intervals. These results demonstrate the value of time-varying higher-order relational modeling for fine-grained manipulation activity recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。