arXiv:2510.14713cs.CVcs.AI2025-10中稿 · AIROV2025

首次评估深度模型在历史影像中的运镜分类表现。

Camera Movement Classification in Historical Footage: A Comparative Study of Deep Video Models

  • 在历史影像数据集上对比五种主流视频模型的运镜识别能力。
  • 最佳模型Video Swin Transformer达80.25%准确率,表现稳定。
  • 适合关注历史视频理解与低质量视频建模的研究者。

运镜传递空间与叙事信息,对理解视频内容至关重要。尽管近期运镜分类(CMC)方法在现代数据集上表现良好,但其在历史影像上的泛化能力仍未知。本文首次系统评估了深度视频CMC模型在档案影片材料上的表现。我们梳理了代表性方法与数据集,指出模型设计与标注定义的差异。在包含专家标注的二战影像的HISTORIAN数据集上,评估了五种标准视频分类模型。表现最佳的Video Swin Transformer达到80.25%准确率,即使在训练数据有限的情况下也表现出强收敛性。研究揭示了现有模型适配低质量视频的挑战与潜力,推动未来结合多模态输入与时序架构的研究。

原文摘要 · Abstract (English)

Camera movement conveys spatial and narrative information essential for understanding video content. While recent camera movement classification (CMC) methods perform well on modern datasets, their generalization to historical footage remains unexplored. This paper presents the first systematic evaluation of deep video CMC models on archival film material. We summarize representative methods and datasets, highlighting differences in model design and label definitions. Five standard video classification models are assessed on the HISTORIAN dataset, which includes expert-annotated World War II footage. The best-performing model, Video Swin Transformer, achieves 80.25% accuracy, showing strong convergence despite limited training data. Our findings highlight the challenges and potential of adapting existing models to low-quality video and motivate future work combining diverse input modalities and temporal architectures.

运镜识别历史影像视频分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。