arXiv:2508.00926cs.LG2025-08

用分段建图方法提升多模态时序数据分类效果

Hybrid Hypergraph Networks for Multimodal Sequence Data Classification

  • 先分段再建图,用超边捕捉模态内结构
  • 在4个数据集上达到当前最优性能
  • 适合处理音频视频等强时序多模态任务

建模时序多模态数据在分类任务中面临挑战,尤其在捕捉长程时序依赖和复杂的跨模态交互方面。以音视频数据为例,其具有严格的时序顺序和多种模态特征。有效利用时序结构对理解模态内动态和模态间相关性至关重要。然而,现有方法通常独立处理各模态并依赖浅层融合策略,忽略了时序依赖,限制了复杂结构关系的表征能力。为此,本文提出混合超图网络(HHN),一种通过‘分段先行、图后构建’策略建模时序多模态数据的新框架。将序列划分为带时间戳的片段作为异构图中的节点,基于最大熵差准则构建超边以捕捉模态内结构,增强节点异质性和结构区分度,再通过超图卷积提取高阶依赖;跨模态连接通过时序对齐与图注意力实现语义融合。HHN在四个多模态数据集上取得最先进的结果,验证了其在复杂分类任务中的有效性。

原文摘要 · Abstract (English)

Modeling temporal multimodal data poses significant challenges in classification tasks, particularly in capturing long-range temporal dependencies and intricate cross-modal interactions. Audiovisual data, as a representative example, is inherently characterized by strict temporal order and diverse modalities. Effectively leveraging the temporal structure is essential for understanding both intra-modal dynamics and inter-modal correlations. However, most existing approaches treat each modality independently and rely on shallow fusion strategies, which overlook temporal dependencies and hinder the model's ability to represent complex structural relationships. To address the limitation, we propose the hybrid hypergraph network (HHN), a novel framework that models temporal multimodal data via a segmentation-first, graph-later strategy. HHN splits sequences into timestamped segments as nodes in a heterogeneous graph. Intra-modal structures are captured via hyperedges guided by a maximum entropy difference criterion, enhancing node heterogeneity and structural discrimination, followed by hypergraph convolution to extract high-order dependencies. Inter-modal links are established through temporal alignment and graph attention for semantic fusion. HHN achieves state-of-the-art (SOTA) results on four multimodal datasets, demonstrating its effectiveness in complex classification tasks.

多模态时序建模超图网络分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。