提出融合音频视频信息的新方法,提升城市监控中事件识别准确率。
Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition

- 用双向交叉注意力融合音视频特征,结合预训练模型提取表示。
- 在AVE数据集上测试准确率达83.85%±1.40%,验证了方法有效性。
- 适合做多模态感知、智能城市监控的研究者参考。
音频-视觉事件识别(AVER)对智能城市监控系统至关重要,需在复杂环境中实现鲁棒的多模态理解。本文提出一种稳定的混合交叉注意力融合框架,用于智能城市环境下的音视频事件识别。该架构结合预训练的Video Masked Autoencoder(VideoMAE)和音频频谱变换器(AST)表示,采用FiLM音视频条件化、双向交叉注意力融合、多模态Transformer编码及模态-时间注意力机制。为提升计算效率与训练稳定性,采用冻结的预训练主干网络与缓存特征提取。在AVE数据集上的大量实验表明,所提框架在多个评估指标下均优于对比的单模态与多模态基线,验证集准确率达到91.74%,测试准确率为83.85%±1.40%(五次独立运行)。结果表明,该混合融合策略能有效捕捉互补的音视频信息,为复杂真实城市监控场景提供稳健的多模态表征学习能力。
原文摘要 · Abstract (English)
Audio-Visual Event Recognition (AVER) is essential for intelligent urban monitoring systems, where robust multimodal understanding of complex environments is required. This paper proposes a stable hybrid cross-attention fusion framework for audio-visual event recognition in smart urban environments. The proposed architecture combines pretrained Video Masked Autoencoder (VideoMAE) and Audio Spectrogram Transformer (AST) representations with FiLM-based audio conditioning, bidirectional cross-attention fusion, multimodal Transformer encoding, and modality-temporal attention. To improve computational efficiency and training stability, frozen pretrained backbones and cached feature extraction are employed. Extensive experiments on the AVE dataset show that the proposed framework achieves the highest average performance among the evaluated unimodal and multimodal baselines across multiple evaluation metrics, obtaining a best validation accuracy of 91.74% and a test accuracy of 83.85 plus/minus 1.40% over five independent runs. The results indicate that the proposed hybrid fusion strategy effectively captures complementary audio-visual information and provides robust multimodal representation learning for challenging realworld urban monitoring scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。