arXiv:2508.17025cs.CVcs.MM2025-08中稿 · IEEE Transactions …被引 4

提出概率时序掩码注意力机制,提升跨视角动作检测泛化能力。

Probabilistic Temporal Masked Attention for Cross-view Online Action Detection

  • 用概率建模压缩视频帧表示,支持跨视角特征提取。
  • 引入GRU-based时序掩码注意力,增强帧间信息交互与自回归分析。
  • 在三个数据集上实现当前最佳性能,适合跨视角场景应用。

作为计算机视觉中视频序列分类的关键任务,在线动作检测(OAD)受到广泛关注。主流OAD模型对不同视频视角敏感,导致在未见源数据上泛化能力差。为此,我们提出一种新型概率时序掩码注意力(PTMA)模型,通过概率建模在跨视角设置下生成视频帧的隐含压缩表示。PTMA融合基于GRU的时序掩码注意力(TMA)单元,利用这些表示有效查询输入视频序列,从而增强信息交互并支持自回归式的帧级视频分析。此外,多视角信息可融入概率建模,促进视图不变特征的提取。在跨主体(cs)、跨视角(cv)及跨主体-视角(csv)三种评估协议下,PTMA在DAHLIA、IKEA ASM和Breakfast数据集上均达到当前最优性能。

原文摘要 · Abstract (English)

As a critical task in video sequence classification within computer vision, Online Action Detection (OAD) has garnered significant attention. The sensitivity of mainstream OAD models to varying video viewpoints often hampers their generalization when confronted with unseen sources. To address this limitation, we propose a novel Probabilistic Temporal Masked Attention (PTMA) model, which leverages probabilistic modeling to derive latent compressed representations of video frames in a cross-view setting. The PTMA model incorporates a GRU-based temporal masked attention (TMA) cell, which leverages these representations to effectively query the input video sequence, thereby enhancing information interaction and facilitating autoregressive frame-level video analysis. Additionally, multi-view information can be integrated into the probabilistic modeling to facilitate the extraction of view-invariant features. Experiments conducted under three evaluation protocols: cross-subject (cs), cross-view (cv), and cross-subject-view (csv) show that PTMA achieves state-of-the-art performance on the DAHLIA, IKEA ASM, and Breakfast datasets.

动作检测跨视角注意力机制视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。