用语义感知轨迹标记提升少样本动作识别性能
Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
- 基于语义与物体尺度自适应采样轨迹点
- 通过方向位移直方图捕捉轨迹内动态关系
- 在6个数据集上达到当前最优效果,适合动作识别研究者
视频理解需有效建模运动与外观信息,尤其在少样本动作识别场景下。尽管点追踪技术已提升识别效果,但仍面临两个核心挑战:如何选取有信息量的追踪点,以及如何有效建模其运动模式。本文提出Trokens,将轨迹点转化为语义感知的关联标记用于动作识别。首先,设计语义感知采样策略,根据物体尺度和语义相关性自适应分布追踪点;其次,构建运动建模框架,通过方向位移直方图(HoD)捕捉轨迹内部动态,并建模轨迹间的相互关系以刻画复杂动作模式。该方法将轨迹标记与语义特征融合,增强外观特征中的运动信息,在六个不同少样本动作识别基准上取得领先表现:Something-Something-V2(全量与小样本划分)、Kinetics、UCF101、HMDB51 和 FineGym。
原文摘要 · Abstract (English)
Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。