arXiv:2602.00701cs.MMcs.CV2026-02

提出线性复杂度的跨模态融合方法,实现高效音频视觉学习。

Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning

  • 用二值化操作实现线性复杂度的跨模态注意力机制。
  • 在CREMA-D、AVE等数据集上超越现有基线模型性能。
  • 适合部署于低功耗边缘设备的音频视觉智能系统。

有效的多模态融合需要既能捕捉复杂跨模态依赖,又具备计算可扩展性的机制。现有音频-视觉融合方法面临根本权衡:基于注意力的方法虽能有效建模跨模态关系,但存在二次计算复杂度,难以支持层次化、多尺度架构;而高效融合策略依赖简单拼接,无法提取互补的跨模态信息。本文提出CMQKA,一种新型跨模态融合机制,通过高效的二值化操作实现线性O(N)复杂度,使此前不可行的层次化融合成为可能。CMQKA采用双向跨模态查询-键注意力,提取互补的时空特征,并利用可学习残差融合保留模态特异性,同时丰富表示。在此基础上,构建了SNNergy框架,采用分层架构,通过逐步降低空间分辨率并提升语义抽象,实现多尺度融合,捕获模态间的局部模式与全局上下文。该框架采用事件驱动的二值脉冲操作,显著提升能效,在挑战性音频-视觉基准数据集CREMA-D、AVE和UrbanSound8K-AV上取得新最佳性能,显著优于现有多模态融合基线。

原文摘要 · Abstract (English)

Effective multimodal fusion requires mechanisms that can capture complex cross-modal dependencies while remaining computationally scalable for real-world deployment. Existing audio-visual fusion approaches face a fundamental trade-off: attention-based methods effectively model cross-modal relationships but incur quadratic computational complexity that prevents hierarchical, multi-scale architectures, while efficient fusion strategies rely on simplistic concatenation that fails to extract complementary cross-modal information. We introduce CMQKA, a novel cross-modal fusion mechanism that achieves linear O(N) complexity through efficient binary operations, enabling scalable hierarchical fusion previously infeasible with conventional attention. CMQKA employs bidirectional cross-modal Query-Key attention to extract complementary spatiotemporal features and uses learnable residual fusion to preserve modality-specific characteristics while enriching representations with cross-modal information. Building upon CMQKA, we present SNNergy, an energy-efficient multimodal fusion framework with a hierarchical architecture that processes inputs through progressively decreasing spatial resolutions and increasing semantic abstraction. This multi-scale fusion capability allows the framework to capture both local patterns and global context across modalities. Implemented with event-driven binary spike operations, SNNergy achieves remarkable energy efficiency while maintaining fusion effectiveness and establishing new state-of-the-art results on challenging audio-visual benchmarks, including CREMA-D, AVE, and UrbanSound8K-AV, significantly outperforming existing multimodal fusion baselines. Our framework advances multimodal fusion by introducing a scalable fusion mechanism that enables hierarchical cross-modal integration with practical energy efficiency for real-world audio-visual intelligence systems.

多模态融合能量效率二值化音频视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。