arXiv:2504.00647cs.CV2025-04被引 1

通过频率解耦提升动作边界精度,解决预训练模型噪声干扰问题

FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection

  • 设计自适应时序解耦机制,过滤无关语义保留原子动作细节
  • 在THUMOS14、HACS、ActivityNet-1.3上超越现有方法,实现最佳定位精度
  • 适合关注视频动作检测边界优化的研究者与应用开发者

时间动作检测旨在定位和分类未剪辑视频中的动作。尽管近期工作聚焦于设计强大的特征处理器以利用预训练表示,但常忽视这些特征中固有的噪声与冗余。大规模预训练视频编码器往往引入背景杂波和无关语义,导致上下文混淆和边界不准确。为此,我们提出一种频率感知的解耦网络,通过滤除预训练模型捕获的噪声语义来提升动作可区分性。具体而言,引入自适应时序解耦方案,在抑制无关信息的同时保留细粒度原子动作细节,生成更具任务特性的表示。同时,通过捕捉时序变化增强帧间建模,更好区分动作与背景冗余。此外,提出长短期类别感知关系网络,联合建模局部过渡与长程依赖,提升定位精度。经过精炼的原子特征与频率引导的动力学信息输入标准检测头,生成精确动作预测。在THUMOS14、HACS和ActivityNet-1.3上的大量实验表明,基于InternVideo2-6B特征的方法在时间动作检测基准上达到当前最优性能。

原文摘要 · Abstract (English)

Temporal action detection aims to locate and classify actions in untrimmed videos. While recent works focus on designing powerful feature processors for pre-trained representations, they often overlook the inherent noise and redundancy within these features. Large-scale pre-trained video encoders tend to introduce background clutter and irrelevant semantics, leading to context confusion and imprecise boundaries. To address this, we propose a frequency-aware decoupling network that improves action discriminability by filtering out noisy semantics captured by pre-trained models. Specifically, we introduce an adaptive temporal decoupling scheme that suppresses irrelevant information while preserving fine-grained atomic action details, yielding more task-specific representations. In addition, we enhance inter-frame modeling by capturing temporal variations to better distinguish actions from background redundancy. Furthermore, we present a long-short-term category-aware relation network that jointly models local transitions and long-range dependencies, improving localization precision. The refined atomic features and frequency-guided dynamics are fed into a standard detection head to produce accurate action predictions. Extensive experiments on THUMOS14, HACS, and ActivityNet-1.3 show that our method, powered by InternVideo2-6B features, achieves state-of-the-art performance on temporal action detection benchmarks.

动作检测边界优化特征解耦视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。