用动态融合提升音视频显著性预测,兼顾精度与效率
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
- 设计动态令牌融合模块,自适应整合视听特征
- 在6个数据集上达到当前最优性能,计算开销低
- 适合做多模态视觉注意力建模的研究者参考
音视频显著性预测旨在通过融合视觉与听觉信息模拟人类视觉注意力,识别视频中的显著区域。尽管纯视觉方法已取得显著进展,但有效融入听觉线索仍面临时空交互复杂、计算成本高的挑战。为此,本文提出动态令牌融合显著性模型(DFTSal),在保证计算效率的同时提升预测精度。该框架采用多尺度视觉编码器,包含两个新模块:可学习令牌增强块(LTEB)自适应加权关键语义令牌,动态可学习令牌融合块(DLTFB)通过移位操作重组并融合特征,有效捕捉长程依赖与细节空间信息。同时,音频分支处理原始音频信号提取有意义的听觉特征。视觉与音频特征通过自适应多模态融合块(AMFB)进行融合,该模块采用局部、全局及自适应三路融合策略实现精确跨模态融合。最终,融合特征经分层多解码器结构生成高精度显著性图。在六个音视频基准数据集上的大量实验表明,DFTSal实现了当前最优性能,且计算效率优异。
原文摘要 · Abstract (English)
Audio-visual saliency prediction aims to mimic human visual attention by identifying salient regions in videos through the integration of both visual and auditory information. Although visual-only approaches have significantly advanced, effectively incorporating auditory cues remains challenging due to complex spatio-temporal interactions and high computational demands. To address these challenges, we propose Dynamic Token Fusion Saliency (DFTSal), a novel audio-visual saliency prediction framework designed to balance accuracy with computational efficiency. Our approach features a multi-scale visual encoder equipped with two novel modules: the Learnable Token Enhancement Block (LTEB), which adaptively weights tokens to emphasize crucial saliency cues, and the Dynamic Learnable Token Fusion Block (DLTFB), which employs a shifting operation to reorganize and merge features, effectively capturing long-range dependencies and detailed spatial information. In parallel, an audio branch processes raw audio signals to extract meaningful auditory features. Both visual and audio features are integrated using our Adaptive Multimodal Fusion Block (AMFB), which employs local, global, and adaptive fusion streams for precise cross-modal fusion. The resulting fused features are processed by a hierarchical multi-decoder structure, producing accurate saliency maps. Extensive evaluations on six audio-visual benchmarks demonstrate that DFTSal achieves SOTA performance while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。