轻量级框架提升多人动作检测的时空稳定性,不依赖复杂模型。
TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection

- 用低抖动检测+高斯特征提取,构建稳定动作管
- 在多体育和UCF101-24数据集上提升4.5~7.1%视频级精度
- 无需光流或大尺度注意力,适合资源受限场景
视频中的时空动作检测需同时定位空间中的人体并识别动作时间边界。现有方法常因帧级检测抖动、碎片化及时间定位不准导致动作管不稳定。许多近期工作引入重型时空变压器或光流管道,带来高计算开销与可扩展性差的问题。本文提出TubeLite,一种轻量级时空动作检测框架,聚焦于稳定动作管构建与边界感知的时间建模。它将每位演员表示为一串关联时间序列的边界框(即动作管),并在空间与语义层面显式施加时间一致性。该方法结合低抖动人体检测、高斯加权特征提取、高效短时时间传播及边界聚焦的时间预测头,避免使用光流与大规模时间注意力。尽管设计紧凑,其仍实现强视频级定位性能:在MultiSports和UCF101-24数据集上,视频[email protected]分别提升4.5和7.1个百分点,参数量与浮点运算量显著低于基于变换器的方案,证明了通过精心设计的轻量级时间建模即可实现高效有效的时空动作检测。
原文摘要 · Abstract (English)
Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge is constructing temporally stable action tubes, as frame-level detectors often suffer from jitter, fragmentation, and imprecise temporal localization. Many recent approaches address this by introducing heavy spatio-temporal transformers or optical-flow-based pipelines, leading to high computational cost and limited scalability. We propose TubeLite, a lightweight framework for spatio-temporal action detection that focuses on stable tube construction and boundary-aware temporal modeling. TubeLite represents each actor as a tube, defined as a sequence of bounding boxes associated with a single actor over time, and explicitly enforces temporal consistency at both the spatial and semantic levels. The method combines low-jitter actor detection, Gaussian-weighted actor feature extraction, efficient short-term temporal propagation, and a boundary-focused temporal prediction head, while avoiding optical flow and large-scale temporal attention. Despite its compact design, TubeLite achieves strong video-level localization performance. It improves [email protected] by 4.5 and 7.1 percentage points over the best compared method on the MultiSports and UCF101-24 datasets, respectively, with substantially fewer parameters and floating-point operations than transformer-based alternatives, demonstrating that effective spatio-temporal action detection can be obtained through principled, lightweight temporal modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。