不训练的视频行为识别增强方法,提升模型鲁棒性。
ReMA: A Training-Free Plug-and-Play Mixing Augmentation for Video Behavior Recognition
- 通过受控混合扩展表示,保持类内稳定性。
- 在多个视频数据集上显著提升泛化能力。
- 无需额外参数,适合部署到现有模型中。
视频行为识别需在复杂的时空变化下保持稳定且具有判别性的表征。然而,现有的视频数据增强策略多为扰动驱动,常引入不可控变化,放大非判别性因素,削弱类内分布结构,并导致不同时间尺度上性能增益不一致。为此,我们提出代表感知混合增强(ReMA),一种即插即用的增强方法,将混合视为受控替换过程,在扩展表示的同时保持类条件稳定性。ReMA包含两个互补机制:首先,表示对齐机制(RAM)在分布对齐约束下进行类内结构化混合,抑制无关类内漂移并增强统计可靠性;其次,动态选择机制(DSM)生成运动感知的时空掩码,定位扰动区域,引导其避开敏感判别区域,促进时序一致性。通过联合控制混合方式与位置,ReMA在无需额外监督或可训练参数的情况下提升表征鲁棒性。在多个视频行为识别基准上的大量实验表明,ReMA在不同时空粒度下均持续提升泛化性和鲁棒性。
原文摘要 · Abstract (English)
Video behavior recognition demands stable and discriminative representations under complex spatiotemporal variations. However, prevailing data augmentation strategies for videos remain largely perturbation-driven, often introducing uncontrolled variations that amplify non-discriminative factors, which finally weaken intra-class distributional structure and representation drift with inconsistent gains across temporal scales. To address these problems, we propose Representation-aware Mixing Augmentation (ReMA), a plug-and-play augmentation strategy that formulates mixing as a controlled replacement process to expand representations while preserving class-conditional stability. ReMA integrates two complementary mechanisms. Firstly, the Representation Alignment Mechanism (RAM) performs structured intra-class mixing under distributional alignment constraints, suppressing irrelevant intra-class drift while enhancing statistical reliability. Then, the Dynamic Selection Mechanism (DSM) generates motion-aware spatiotemporal masks to localize perturbations, guiding them away from discrimination-sensitive regions and promoting temporal coherence. By jointly controlling how and where mixing is applied, ReMA improves representation robustness without additional supervision or trainable parameters. Extensive experiments on diverse video behavior benchmarks demonstrate that ReMA consistently enhances generalization and robustness across different spatiotemporal granularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。