用强化学习找视频模型薄弱点,攻击更隐蔽、更高效。
Robustness Evaluation for Video Models with Reinforcement Learning
- 多智能体强化学习协同定位视频时空敏感区域
- 在HMDB-51和UCF-101上优于现有方法,减少扰动与查询次数
- 支持自定义扰动类型,适配实际应用场景的鲁棒性评估
视频分类模型的鲁棒性评估极具挑战性,尤其相比图像模型。由于时间维度的增加,其复杂度和计算成本显著上升。关键难点在于如何以最小扰动诱导误分类。本文提出一种基于多智能体强化学习(空间与时间)的方法,协同学习识别视频中敏感的空间与时间区域。智能体在生成扰动时考虑时间连贯性,使攻击更有效且视觉上难以察觉。方法在Lp度量和平均查询次数上超越当前最优方案,并支持自定义扰动类型,使鲁棒性评估更贴近实际使用场景。我们在两个主流数据集HMDB-51和UCF-101上对4种主流视频动作识别模型进行了广泛评估。
原文摘要 · Abstract (English)
Evaluating the robustness of Video classification models is very challenging, specifically when compared to image-based models. With their increased temporal dimension, there is a significant increase in complexity and computational cost. One of the key challenges is to keep the perturbations to a minimum to induce misclassification. In this work, we propose a multi-agent reinforcement learning approach (spatial and temporal) that cooperatively learns to identify the given video's sensitive spatial and temporal regions. The agents consider temporal coherence in generating fine perturbations, leading to a more effective and visually imperceptible attack. Our method outperforms the state-of-the-art solutions on the Lp metric and the average queries. Our method enables custom distortion types, making the robustness evaluation more relevant to the use case. We extensively evaluate 4 popular models for video action recognition on two popular datasets, HMDB-51 and UCF-101.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。