用跨注意力GRU融合多模态信息,提升细粒度视频理解能力
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
- 基于GRU和跨模态注意力融合视频、图像、文本特征
- 在两个数据集上显著优于单模态基线,提升效果明显
- 适合需要多模态分析的视频理解任务研究者
细粒度视频分类需理解复杂时空与语义线索,常超出单一模态能力。本文提出一种多模态框架,通过基于GRU的序列编码器与跨模态注意力机制融合视频、图像和文本表示。模型采用分类或回归损失训练,并结合特征级增强与自编码正则化。在真实暴力检测(DVD数据集)和情绪估测(Aff-Wild2数据集)两个挑战性基准上验证,结果表明该融合策略显著优于单模态基线,跨注意力与特征增强对性能与鲁棒性贡献显著。
原文摘要 · Abstract (English)
Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text representations using GRU-based sequence encoders and cross-modal attention mechanisms. The model is trained using a combination of classification or regression loss, depending on the task, and is further regularized through feature-level augmentation and autoencoding techniques. To evaluate the generality of our framework, we conduct experiments on two challenging benchmarks: the DVD dataset for real-world violence detection and the Aff-Wild2 dataset for valence-arousal estimation. Our results demonstrate that the proposed fusion strategy significantly outperforms unimodal baselines, with cross-attention and feature augmentation contributing notably to robustness and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。