arXiv:2506.20342cs.CVcs.AI2025-06IJCV被引 9

通过生成缺失特征提升动作识别精度,无需增加计算开销。

Feature Hallucination for Self-supervised Action Recognition

  • 用视觉帧联合预测动作与辅助特征,测试时生成缺失线索。
  • 在Kinetics-400等数据集上达顶尖性能,提升细粒度动作识别能力。
  • 适合研究多模态自监督学习与视频理解的开发者参考。

理解视频中的人类动作不仅依赖原始像素分析,还需高层次语义推理和有效融合多模态特征。我们提出一种深度可迁移动作识别框架,通过从RGB视频帧中联合预测动作概念与辅助特征,提升识别准确率。测试时,幻觉流推断缺失线索,丰富特征表示而无需增加计算开销。为聚焦超越原始像素的动作相关区域,引入两种新型领域特定描述符:物体检测特征(ODF)整合多个检测器输出以捕捉上下文线索,显著性检测特征(SDF)突出对动作识别关键的空间与强度模式。该框架可无缝集成光学流、改进密集轨迹、骨骼数据及音频线索等辅助模态。兼容当前主流架构如I3D、AssembleNet、Video Transformer Network、FASTER,以及VideoMAE V2与InternVideo2等新模型。为处理辅助特征的不确定性,引入认知不确定性建模与鲁棒损失函数以抑制特征噪声。在Kinetics-400、Kinetics-600和Something-Something V2等多个基准上达到最先进水平,验证其捕捉细微动作动态的有效性。

原文摘要 · Abstract (English)

Understanding human actions in videos requires more than raw pixel analysis; it relies on high-level semantic reasoning and effective integration of multimodal features. We propose a deep translational action recognition framework that enhances recognition accuracy by jointly predicting action concepts and auxiliary features from RGB video frames. At test time, hallucination streams infer missing cues, enriching feature representations without increasing computational overhead. To focus on action-relevant regions beyond raw pixels, we introduce two novel domain-specific descriptors. Object Detection Features (ODF) aggregate outputs from multiple object detectors to capture contextual cues, while Saliency Detection Features (SDF) highlight spatial and intensity patterns crucial for action recognition. Our framework seamlessly integrates these descriptors with auxiliary modalities such as optical flow, Improved Dense Trajectories, skeleton data, and audio cues. It remains compatible with state-of-the-art architectures, including I3D, AssembleNet, Video Transformer Network, FASTER, and recent models like VideoMAE V2 and InternVideo2. To handle uncertainty in auxiliary features, we incorporate aleatoric uncertainty modeling in the hallucination step and introduce a robust loss function to mitigate feature noise. Our multimodal self-supervised action recognition framework achieves state-of-the-art performance on multiple benchmarks, including Kinetics-400, Kinetics-600, and Something-Something V2, demonstrating its effectiveness in capturing fine-grained action dynamics.

动作识别自监督学习多模态融合特征生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。