arXiv:2510.16444cs.CVcs.MM2025-10被引 5

用多轨迹Mamba提升跨模态对齐,精准定位视频中特定人的细粒度动作。

RefAtomNet++: Advancing Referring Atomic Video Action Recognition using Semantic Retrieval based Multi-Trajectory Mamba

  • 通过多层级语义对齐注意力与多轨迹Mamba建模,增强跨模态信息聚合。
  • 在超290万帧、7.5万+人物的RefAVA++数据集上实现新最优性能。
  • 适合关注视频中个体细粒度动作理解与自然语言引导分析的研究者。

指称原子级视频动作识别(RAVAR)旨在根据自然语言描述,识别特定目标人物的细粒度原子动作。与传统动作识别和检测任务不同,RAVAR强调精确的语言引导动作理解,尤其适用于复杂多人场景下的交互式动作分析。本文将先前的RefAVA数据集扩展为RefAVA++,包含超过290万帧和75.1万标注人物。我们基于多个相关领域基线(包括原子动作定位、视频问答和文本-视频检索)及早期模型RefAtomNet对该数据集进行基准测试。尽管RefAtomNet通过代理注意力突出显著特征,但其跨模态信息对齐与检索能力仍有限,导致目标人物定位与细粒度动作预测表现不佳。为此,本文提出RefAtomNet++,一种新框架,通过多层级语义对齐交叉注意力机制结合局部关键词、场景属性和整体句子层面的多轨迹Mamba建模,实现更优的跨模态标记聚合。具体而言,在局部关键词和场景属性层面,动态选择每时刻最近的视觉空间标记构建扫描轨迹。同时设计多层级语义对齐交叉注意力策略,有效聚合不同语义层级的空间与时间标记。实验表明,RefAtomNet++达到新的最先进水平。数据集与代码已开源:https://github.com/KPeng9510/refAVA2。

原文摘要 · Abstract (English)

Referring Atomic Video Action Recognition (RAVAR) aims to recognize fine-grained, atomic-level actions of a specific person of interest conditioned on natural language descriptions. Distinct from conventional action recognition and detection tasks, RAVAR emphasizes precise language-guided action understanding, which is particularly critical for interactive human action analysis in complex multi-person scenarios. In this work, we extend our previously introduced RefAVA dataset to RefAVA++, which comprises >2.9 million frames and >75.1k annotated persons in total. We benchmark this dataset using baselines from multiple related domains, including atomic action localization, video question answering, and text-video retrieval, as well as our earlier model, RefAtomNet. Although RefAtomNet surpasses other baselines by incorporating agent attention to highlight salient features, its ability to align and retrieve cross-modal information remains limited, leading to suboptimal performance in localizing the target person and predicting fine-grained actions. To overcome the aforementioned limitations, we introduce RefAtomNet++, a novel framework that advances cross-modal token aggregation through a multi-hierarchical semantic-aligned cross-attention mechanism combined with multi-trajectory Mamba modeling at the partial-keyword, scene-attribute, and holistic-sentence levels. In particular, scanning trajectories are constructed by dynamically selecting the nearest visual spatial tokens at each timestep for both partial-keyword and scene-attribute levels. Moreover, we design a multi-hierarchical semantic-aligned cross-attention strategy, enabling more effective aggregation of spatial and temporal tokens across different semantic hierarchies. Experiments show that RefAtomNet++ establishes new state-of-the-art results. The dataset and code are released at https://github.com/KPeng9510/refAVA2.

动作识别跨模态对齐Mamba细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。