arXiv:2511.21202cs.CV2025-11

通过追踪局部区域动态,提升细粒度视频动作识别精度

Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition

  • 用查询-响应机制定位并跟踪视频中关键局部区域的动态变化
  • 在多个基准数据集上超越现有最优方法,显著提升细粒度动作区分能力
  • 适合需要精准捕捉细微动作差异的研究者与应用开发者

细粒度动作识别(FGAR)旨在区分类别间细微差异。现有方法多关注粗粒度运动模式,难以捕捉随时间演变的局部细节。本文提出动作区域追踪(ART)框架,利用查询-响应机制发现并追踪具有辨别性的局部区域动态,实现相似动作的有效区分。具体地,设计区域语义激活模块,以判别性文本约束语义作为查询,捕获每帧中最具动作相关性的区域响应,促进时空维度间特征交互。捕获的区域响应被组织为动作轨迹片段,通过跨帧连贯关联刻画基于区域的动作演化过程。文本约束查询由视觉语言模型中的语言分支提取的动作标签描述生成,编码细微语义信息。设计多层次轨迹对比约束,在空间与时间层面优化轨迹,增强帧内区分力与帧间关联性。此外,引入任务特异性微调机制,在保留VLM语义的前提下优化任务适配性。在多个主流动作识别基准上的实验表明,该方法显著优于当前最先进基线。

原文摘要 · Abstract (English)

Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the Action-Region Tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling effective distinction of similar actions. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are organized into action tracklets, which characterize region-based action dynamics by linking related responses across video frames in a coherent sequence. The text-constrained queries encode nuanced semantic representations derived from textual descriptions of action labels extracted by language branches within Visual Language Models (VLMs). To optimize the action tracklets, we design a multi-level tracklet contrastive constraint among region responses at spatial and temporal levels, enabling effective discrimination within each frame and correlation between adjacent frames. Additionally, a task-specific fine-tuning mechanism refines textual semantics such that semantic representations encoded by VLMs are preserved while optimized for task preferences. Comprehensive experiments on widely used action recognition benchmarks demonstrate the superiority to previous state-of-the-art baselines.

细粒度识别视频分析动作追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。