arXiv:2409.10641cs.CV2024-09ECCV

用分层嵌入技术加速视频标注,10倍减少点击次数。

HAVANA: Hierarchical stochastic neighbor embedding for Accelerated Video ANnotAtions

  • 通过分层随机邻域嵌入构建视频特征多尺度表示
  • 标注12小时视频点击量减少超10倍
  • 适合大规模视频数据标注团队使用

视频标注是计算机视觉研究与应用中的关键且耗时的任务。本文提出一种新颖的标注流程,利用预提取特征和降维技术加速时间序列视频标注。该方法采用分层随机邻域嵌入(HSNE)创建视频特征的多尺度表示,使标注者能高效探索并标记大规模视频数据集。实验表明,相比传统线性方法,本方法显著降低标注开销,在多个数据集上实现超过10倍的点击数减少,用于标注超过12小时视频。我们还研究了不同数据集下HSNE参数的最优配置。该工作为视频理解时代的大规模视频标注提供了有前景的方向。

原文摘要 · Abstract (English)

Video annotation is a critical and time-consuming task in computer vision research and applications. This paper presents a novel annotation pipeline that uses pre-extracted features and dimensionality reduction to accelerate the temporal video annotation process. Our approach uses Hierarchical Stochastic Neighbor Embedding (HSNE) to create a multi-scale representation of video features, allowing annotators to efficiently explore and label large video datasets. We demonstrate significant improvements in annotation effort compared to traditional linear methods, achieving more than a 10x reduction in clicks required for annotating over 12 hours of video. Our experiments on multiple datasets show the effectiveness and robustness of our pipeline across various scenarios. Moreover, we investigate the optimal configuration of HSNE parameters for different datasets. Our work provides a promising direction for scaling up video annotation efforts in the era of video understanding.

视频标注降维HSNE效率提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。