构建大规模遮挡视频数据集,提升动作识别在遮挡下的鲁棒性。
OccludeNet: A Causal Journey into Mixed-View Actor-Centric Video Action Recognition under Occlusions
- 基于因果模型设计遮挡下动作识别新方法,利用反事实推理增强关键主体信息。
- 实测显示部分身体可见的动作在遮挡后准确率下降超30%。
- 适合关注视觉感知鲁棒性、视频理解与因果推理的研究者。
现有动作识别数据集缺乏遮挡场景,制约模型鲁棒性提升。我们构建了OccludeNet,一个大规模真实与合成遮挡视频数据集,涵盖动态遮挡、静态遮挡及多视角交互遮挡,覆盖多种自然场景。分析表明,与场景无关且部分肢体可见的动作在遮挡下准确率下降更显著。为突破现有遮挡感知方法局限,我们提出结构化因果模型,并引入因果动作识别(CAR)方法,通过后门调整与反事实推理强化关键演员信息,提升模型对遮挡的鲁棒性。我们希望OccludeNet能推动遮挡场景中因果关系研究,促进类别间关系再思考,实现长期性能提升。代码与数据已开源。
原文摘要 · Abstract (English)
The lack of occlusion data in common action recognition video datasets limits model robustness and hinders consistent performance gains. We build OccludeNet, a large-scale occluded video dataset including both real and synthetic occlusion scenes in different natural settings. OccludeNet includes dynamic occlusion, static occlusion, and multi-view interactive occlusion, addressing gaps in current datasets. Our analysis shows occlusion affects action classes differently: actions with low scene relevance and partial body visibility see larger drops in accuracy. To overcome the limits of existing occlusion-aware methods, we propose a structural causal model for occluded scenes and introduce the Causal Action Recognition (CAR) method, which uses backdoor adjustment and counterfactual reasoning. This approach strengthens key actor information and improves model robustness to occlusion. We hope the challenges of OccludeNet will encourage more study of causal links in occluded scenes and lead to a fresh look at class relations, ultimately leading to lasting performance improvements. Our code and data is availibale at: https://github.com/The-Martyr/OccludeNet-Dataset
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。