提出首个雾天动作识别数据集与模型,让算法在大雾中仍能准确识行动作。
Seeing Through Fog: Towards Fog-Invariant Action Recognition

- 构建双流CLIP模型,利用清晰视频引导雾中视频学不变特征。
- 在近万段视频的FogAct数据集上实现媲美顶尖方法的性能。
- 适合做户外视觉系统、自动驾驶等真实场景下动作识别研究者。
雾天是现实应用中的常见挑战,但现有动作识别方法通常假设理想天气和高质量视频输入。雾霾导致能见度不稳、对比度下降,干扰语义线索提取,严重影响当前方法表现。本文通过两种策略缓解雾天动作识别难题:首先,提出FogAct,首个用于雾天动作识别的基准数据集,由双目相机系统采集的成对清晰与雾化视频构成,涵盖10个场景、55类动作,共近10,000段视频片段;其次,提出FogNet,一种双流CLIP模型,可挖掘隐藏在退化视频后的雾不变语义信息。FogNet借助清晰视频引导,学习雾化视频的鲁棒表征,有效捕捉清晰与雾化视频间的共享结构与运动线索。在FogAct及三个其他主流数据集上的大量实验表明,该方法性能媲美当前最先进(SOTA)水平。FogAct与FogNet已公开于项目主页。
原文摘要 · Abstract (English)
Foggy conditions are commonly encountered in real-world applications; however, existing action recognition approaches typically assume favorable weather and high-quality video inputs. On foggy days, unpredictable visibility degradation and reduced contrast obstruct the extraction of semantic cues, posing significant challenges for current action recognition methods. In this paper, we mitigate the issues faced in action recognition under foggy conditions by employing two strategies. First, we present FogAct, the first benchmark dataset for foggy action recognition, consisting of paired clean and foggy videos captured with a stereo camera system. The dataset spans 10 scenes and 55 action categories, comprising nearly 10,000 video clips. Second, we propose FogNet, a two-stream CLIP model that discovers fog-invariant semantic information hidden behind the degraded videos. FogNet learns robust representations of foggy videos with guidance from clean videos, effectively capturing shared structural and motion cues between clean and foggy videos. Extensive experiments on FogAct and three other popular datasets demonstrate that our method achieves competitive performance compared with state-of-the-art (SOTA) approaches. Our FogAct and FogNet are given in our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。