arXiv:2506.00358cs.SDcs.AI2025-06NeurIPS被引 3

构建多模态联合干扰测试集,评估音视频模型真实场景下的鲁棒性

$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time

  • 设计四组双模态共现干扰数据集,模拟真实世界中音视频同时失真
  • 顶尖音视频模型在严重干扰下准确率显著下降,现有自适应方法提升有限
  • 提出轻量级跨模态融合策略,通过抑制高熵样本提升模型稳定性

尽管近期音视频模型表现优异,但其在测试阶段面对分布偏移的鲁棒性仍不明确。现有基准多聚焦单一模态,难以全面评估音视频模型。针对音视频模态可能同时发生相关性偏移的真实场景,我们提出 $ exttt{AVROBUSTBENCH}$,一个涵盖 $ exttt{AUDIOSET-2C}$、$ exttt{VGGSOUND-2C}$、$ exttt{KINETICS-2C}$ 与 $ exttt{EPICKITCHENS-2C}$ 的综合性基准,每项均包含 75 种共现且相关的双模态干扰。实验表明,先进监督与自监督模型在干扰强度增加时性能持续下降。在线测试时自适应(TTA)方法在 $ exttt{VGGSOUND-2C}$ 与 $ exttt{KINETICS-2C}$ 上改善微弱。我们进一步提出 $ exttt{AV2C}$,一种通过惩罚高熵样本实现动态跨模态融合的轻量级 TTA 方法,在 $ exttt{VGGSOUND-2C}$ 上取得性能提升。期待 $ exttt{AVROBUSTBENCH}$ 推动更有效鲁棒的音视频测试时自适应方法发展。

原文摘要 · Abstract (English)

While recent audio-visual models have demonstrated impressive performance, their robustness to distributional shifts at test-time remains not fully understood. Existing robustness benchmarks mainly focus on single modalities, making them insufficient for thoroughly assessing the robustness of audio-visual models. Motivated by real-world scenarios where shifts can occur $\textit{simultaneously}$ in both audio and visual modalities, we introduce $\texttt{AVROBUSTBENCH}$, a comprehensive benchmark designed to evaluate the test-time robustness of audio-visual recognition models. $\texttt{AVROBUSTBENCH}$ comprises four audio-visual benchmark datasets, $\texttt{AUDIOSET-2C}$, $\texttt{VGGSOUND-2C}$, $\texttt{KINETICS-2C}$, and $\texttt{EPICKITCHENS-2C}$, each incorporating 75 bimodal audio-visual corruptions that are $\textit{co-occurring}$ and $\textit{correlated}$. Through extensive evaluations, we observe that state-of-the-art supervised and self-supervised audio-visual models exhibit declining robustness as corruption severity increases. Furthermore, online test-time adaptation (TTA) methods, on $\texttt{VGGSOUND-2C}$ and $\texttt{KINETICS-2C}$, offer minimal improvements in performance under bimodal corruptions. We further propose $\texttt{AV2C}$, a simple TTA approach enabling on-the-fly cross-modal fusion by penalizing high-entropy samples, which achieves improvements on $\texttt{VGGSOUND-2C}$. We hope that $\texttt{AVROBUSTBENCH}$ will steer the development of more effective and robust audio-visual TTA approaches. Our code is available $\href{https://github.com/sarthaxxxxx/AV-C-Robustness-Benchmark}{here}$.

音视频融合测试时自适应鲁棒性评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。