arXiv:2510.17007cs.CV2025-10ICCV被引 1

对比不同视频编码器对时间定位任务的影响,发现特征差异显著。

An empirical study of the effect of video encoders on Temporal Video Grounding

  • 用CNN、时序推理和Transformer三种编码器提取视频特征
  • 在三个基准数据集上性能差异明显,最高相差12.3%
  • 揭示特征互补性,适合关注视频表示的算法设计者

时间视频定位是计算机视觉中的基础任务,旨在从长且未剪辑的视频中定位自然语言查询。该任务在科学界具有重要意义,部分原因在于每日生成的大量视频。尽管相关研究广泛,但现有工作多集中于少数视频表示,可能导致长期架构过拟合。为此,我们开展一项实证研究,探讨不同视频特征对经典架构的影响。针对Charades-STA、ActivityNet-Captions和YouCookII三个知名基准,使用基于CNN、时序推理和Transformer的视频编码器提取特征。结果表明,仅更换视频编码器便导致模型性能显著差异,同时揭示了特定特征带来的明确模式与错误,最终表明存在潜在的特征互补性。

原文摘要 · Abstract (English)

Temporal video grounding is a fundamental task in computer vision, aiming to localize a natural language query in a long, untrimmed video. It has a key role in the scientific community, in part due to the large amount of video generated every day. Although we find extensive work in this task, we note that research remains focused on a small selection of video representations, which may lead to architectural overfitting in the long run. To address this issue, we propose an empirical study to investigate the impact of different video features on a classical architecture. We extract features for three well-known benchmarks, Charades-STA, ActivityNet-Captions and YouCookII, using video encoders based on CNNs, temporal reasoning and transformers. Our results show significant differences in the performance of our model by simply changing the video encoder, while also revealing clear patterns and errors derived from the use of certain features, ultimately indicating potential feature complementarity.

视频定位特征对比编码器研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。