arXiv:2504.00149cs.CVcs.AI2025-04

解决标注时间错位问题,提升动作定位精度

Towards Precise Action Spotting: Addressing Temporal Misalignment in Labels with Dynamic Label Assignment

  • 提出动态标签分配策略,允许预测时间偏移
  • 在视觉显著事件上实现当前最佳性能
  • 适合需要高时序精度的动作识别任务

精确动作定位因广泛应用而受到关注。尽管现有方法通过精心设计的模型架构取得了显著性能,却忽视了一个关键挑战:真实标签中的时间错位问题。这种错位源于标注时帧与实际事件时间不匹配,常由人工标注误差或相邻帧事件边界判断困难导致。为此,本文提出一种新型动态标签分配策略,使训练时预测可偏离真实动作时间,确保一致的动作定位。该方法将最小代价匹配思想从空间目标检测拓展至时间域,基于预测类别得分和时间偏移计算匹配代价,动态为最可能的预测分配标签,有效缓解标签时间错位带来的负面影响。大量实验表明,该方法在事件视觉显著且标签错位普遍的场景下达到领先性能。

原文摘要 · Abstract (English)

Precise action spotting has attracted considerable attention due to its promising applications. While existing methods achieve substantial performance by employing well-designed model architecture, they overlook a significant challenge: the temporal misalignment inherent in ground-truth labels. This misalignment arises when frames labeled as containing events do not align accurately with the actual event times, often as a result of human annotation errors or the inherent difficulties in precisely identifying event boundaries across neighboring frames. To tackle this issue, we propose a novel dynamic label assignment strategy that allows predictions to have temporal offsets from ground-truth action times during training, ensuring consistent event spotting. Our method extends the concept of minimum-cost matching, which is utilized in the spatial domain for object detection, to the temporal domain. By calculating matching costs based on predicted action class scores and temporal offsets, our method dynamically assigns labels to the most likely predictions, even when the predicted times of these predictions deviate from ground-truth times, alleviating the negative effects of temporal misalignment in labels. We conduct extensive experiments and demonstrate that our method achieves state-of-the-art performance, particularly in conditions where events are visually distinct and temporal misalignment in labels is common.

动作定位时间对齐动态分配视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。