arXiv:2604.03329cs.CVcs.AI2026-04

视觉动态调控音频处理,让模型在关键时刻依赖音频。

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

论文配图:AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection
图 1 · 摘自论文原文
  • 视觉信息实时调节音频编码器的时序行为,决定何时用音频。
  • 在缺失或噪声音频下仍保持88.59%准确率,优于固定路由方法。
  • 适合做音视频融合的安防、内容审核系统,尤其音频不可靠时。

从视频中自动检测暴力行为极具挑战,因暴力动作可能距离远、被遮挡或部分可见。音频可提供仅凭视觉难以识别的补充线索,但音频本身可能缺失、被替换或被环境噪音掩盖。因此核心挑战不在于是否引入音频,而在于如何根据视觉场景动态调整对音频的依赖程度。我们提出AViS-Mamba,一种基于Mamba的音视频架构,其中视觉流直接控制音频流的行为。在音频编码器每一层,一个紧凑的视觉表征生成调制向量,共同调节编码器内部的时序操作,并通过路由门控制视觉干预强度。不同于提取特征后融合或重加权,视觉上下文直接塑造音频编码器的时序动态。我们还提出自适应的AV-InfoNCE对比目标,学习平衡音频到视频与视频到音频的对齐方向,而非均匀加权。在音频有效的NTU-CCTV和DVD数据集上,AViS-Mamba达到88.59%和75.74%的准确率,为当前最优。实验表明,自适应视觉调制始终优于固定路由,在音频退化或缺失条件下性能更优。层间分析显示,模型在不同网络深度选择性地调节音频流,而非采用单一全局策略。

原文摘要 · Abstract (English)

Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Audio can provide complementary evidence for violent events that are difficult to recognize from visual information alone. However, audio itself may be absent, dubbed, or dominated by environmental noise, making the central challenge not whether to incorporate audio but how to adapt reliance on it according to the visual scene. We introduce \emph{AViS-Mamba}, an audiovisual Mamba-based architecture in which the visual stream directly governs the behavior of the audio stream. At each layer of the audio encoder, a compact visual representation produces a modulation vector that conditions the encoder's internal temporal operators together with a routing gate that regulates the strength of this visual intervention. Rather than fusing or reweighting features after they have been extracted, visual context directly shapes the temporal dynamics of the audio encoder. We further propose Adaptive AV-InfoNCE, a contrastive objective that learns to balance the audio-to-video and video-to-audio alignment directions rather than weighting them uniformly. On the audio-valid NTU-CCTV and DVD benchmarks, AViS-Mamba establishes state-of-the-art results, attaining 88.59% and 75.74% accuracy. We demonstrate that adaptive visual conditioning consistently outperforms fixed routing and improves performance under degraded and missing-audio conditions. Layer-wise analysis further reveals that the model adapts the audio stream selectively across network depth rather than applying a single global routing policy.

暴力检测音视频融合自适应机制Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。