arXiv:2602.01257cs.CV2026-02被引 1

用文本优化视频动作定位,提升准确率。

Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment

  • 引入文本精炼与多模态对齐模块,融合语义信息
  • 在5个基准上超越现有方法,精度显著提升
  • 仅需单张3090显卡,适合实际部署

点标注的时序动作定位近年受到关注,因其在标注成本与定位精度间取得良好平衡。然而,现有方法仅依赖视觉特征,忽视了文本侧的语义信息。为此,本文提出文本精炼与对齐(TRA)框架,利用预训练多模态模型生成的视频帧描述,增强视觉特征的语义表达。设计两个新模块:基于点的文本精炼模块(PTR)和基于点的多模态对齐模块(PMA)。首先用预训练模型生成帧级描述;随后,PTR结合点标注与多个预训练模型精炼文本;PMA将视觉与语言特征投影至统一语义空间,并通过点级别对比学习缩小模态差距。最终,融合后的多模态特征输入动作检测器实现精准定位。在五个主流基准上的实验表明,本框架优于多个先进方法。计算开销分析显示,可在单张24 GB RTX 3090 GPU上运行,具备实用性与可扩展性。

原文摘要 · Abstract (English)

Recently, point-supervised temporal action localization has gained significant attention for its effective balance between labeling costs and localization accuracy. However, current methods only consider features from visual inputs, neglecting helpful semantic information from the text side. To address this issue, we propose a Text Refinement and Alignment (TRA) framework that effectively utilizes textual features from visual descriptions to complement the visual features as they are semantically rich. This is achieved by designing two new modules for the original point-supervised framework: a Point-based Text Refinement module (PTR) and a Point-based Multimodal Alignment module (PMA). Specifically, we first generate descriptions for video frames using a pre-trained multimodal model. Next, PTR refines the initial descriptions by leveraging point annotations together with multiple pre-trained models. PMA then projects all features into a unified semantic space and leverages a point-level multimodal feature contrastive learning to reduce the gap between visual and linguistic modalities. Last, the enhanced multi-modal features are fed into the action detector for precise localization. Extensive experimental results on five widely used benchmarks demonstrate the favorable performance of our proposed framework compared to several state-of-the-art methods. Moreover, our computational overhead analysis shows that the framework can run on a single 24 GB RTX 3090 GPU, indicating its practicality and scalability.

动作定位多模态文本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。