arXiv:2509.12554cs.CV2025-09

用偏微分方程建模图信息传播,提升人-物交互检测精度

Multimodal Graph Network Modeling for Human-Object Interaction Detection with PDE Graph Diffusion

  • 构建四阶段多模态图网络,用PDE方程实现可解释的信息扩散
  • 在HICO-DET和V-COCO上达最新水平,稀有类别表现更均衡
  • 适合关注图神经网络可解释性与视觉语言模型噪声抑制的研究者

现有基于图神经网络的交互检测方法依赖简单MLP融合实例特征并传播信息,但该机制高度经验化,缺乏定向传播设计。为此,我们提出多模态图网络建模(MGNM)用于交互检测,引入偏微分方程(PDE)图扩散机制。首先设计四阶段多模态图结构显式建模交互任务;其次提出新型PDE扩散机制,利用多模态特征通过白盒PDE方程传播信息;此外设计变分信息压缩(VIS)机制,进一步优化CLIP提取的多模态特征,缓解预训练视觉-语言模型固有的噪声影响。大量实验表明,所提MGNM在两个主流基准HICO-DET和V-COCO上均达到领先性能;当与更先进的目标检测器结合时,仍保持对稀有与非稀有类别的良好平衡。

原文摘要 · Abstract (English)

Existing GNN-based Human-Object Interaction (HOI) detection methods rely on simple MLPs to fuse instance features and propagate information. However, this mechanism is largely empirical and lack of targeted information propagation process. To address this problem, we propose Multimodal Graph Network Modeling (MGNM) for HOI detection with Partial Differential Equation (PDE) graph diffusion. Specifically, we first design a multimodal graph network framework that explicitly models the HOI detection task within a four-stage graph structure. Next, we propose a novel PDE diffusion mechanism to facilitate information propagation within this graph. This mechanism leverages multimodal features to propaganda information via a white-box PDE diffusion equation. Furthermore, we design a variational information squeezing (VIS) mechanism to further refine the multimodal features extracted from CLIP, thereby mitigating the impact of noise inherent in pretrained Vision-Language Models. Extensive experiments demonstrate that our MGNM achieves state-of-the-art performance on two widely used benchmarks: HICO-DET and V-COCO. Moreover, when integrated with a more advanced object detector, our method yields significant performance gains while maintaining an effective balance between rare and non-rare categories.

图神经网络交互检测PDE扩散多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。