arXiv:2506.03440cs.CV2025-06中稿 · Expert Systems wit…被引 10

通过几何与视觉融合图网络,提升视频中多人物-物体交互识别精度

Geometric Visual Fusion Graph Neural Networks for Multi-Person Human-Object Interaction Recognition in Videos

  • 先构建实体特异性表示,再逐层融合几何与视觉特征
  • 在多场景下实现当前最优性能,尤其擅长处理部分参与和并发动作
  • 新数据集MPHOI-120支持真实复杂交互场景评估

视频中的人物-物体交互(HOI)识别需要理解随时间演化的视觉模式与几何关系。视觉特征捕捉外观上下文,几何特征提供结构模式,但如何在不损失各自优势的前提下有效融合多模态特征仍具挑战。我们发现,在建模交互前建立稳健的、针对实体的表示有助于保持各模态特性,因此提出自底向上的融合策略。为此,我们设计了几何视觉融合图神经网络(GeoVis-GNN),结合双注意力特征融合与相互依赖的实体图学习,逐步从实体特异性表示构建高层次交互理解。为推动HOI识别向真实场景发展,我们引入并发部分交互数据集MPHOI-120,其捕捉涉及并发动作与部分参与的动态多人交互。该数据集应对复杂人-物动态与相互遮挡等挑战。大量实验表明,该方法在多种HOI场景中表现优异,包括两人交互、单人活动、双手操作及复杂并发部分交互,达到当前最优性能。

原文摘要 · Abstract (English)

Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture appearance context, while geometric features provide structural patterns. Effectively fusing these multimodal features without compromising their unique characteristics remains challenging. We observe that establishing robust, entity-specific representations before modeling interactions helps preserve the strengths of each modality. Therefore, we hypothesize that a bottom-up approach is crucial for effective multimodal fusion. Following this insight, we propose the Geometric Visual Fusion Graph Neural Network (GeoVis-GNN), which uses dual-attention feature fusion combined with interdependent entity graph learning. It progressively builds from entity-specific representations toward high-level interaction understanding. To advance HOI recognition to real-world scenarios, we introduce the Concurrent Partial Interaction Dataset (MPHOI-120). It captures dynamic multi-person interactions involving concurrent actions and partial engagement. This dataset helps address challenges like complex human-object dynamics and mutual occlusions. Extensive experiments demonstrate the effectiveness of our method across various HOI scenarios. These scenarios include two-person interactions, single-person activities, bimanual manipulations, and complex concurrent partial interactions. Our method achieves state-of-the-art performance.

视频理解图神经网络交互识别多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。