arXiv:2510.15725cs.CVcs.AI2025-10中稿 · ACMMM2025 SUMAC

为老电影运动分析设计轻量级编码器,提升模型在模糊低质画面下的准确率。

DGME-T: Directional Grid Motion Encoding for Transformer-Based Historical Camera Movement Classification

  • 用光流生成方向网格编码,通过可学习融合层注入Transformer模型
  • 在现代视频上准确率提升至86.14%,二战影片也从83.43%增至84.62%
  • 适合研究历史影像分析、老旧视频处理的开发者和研究人员

基于现代高质量视频训练的相机运动分类(CMC)模型在档案电影上表现下降,因噪声、丢帧与低对比度干扰运动线索。本文构建统一基准,整合两个现代数据集并重构HISTORIAN为五个均衡类别。在此基础上提出DGME-T,作为Video Swin Transformer的轻量扩展,通过可学习且归一化的后期融合层注入由光流生成的方向性网格运动编码。DGME-T将骨干网络在现代视频上的顶1准确率从81.78%提升至86.14%,宏平均F1从82.08%升至87.81%;在高难度二战影片上,准确率从83.43%提升至84.62%,宏平均F1从81.72%增至82.63%。跨域实验表明,中间阶段在现代数据上微调可使历史影像性能提升超5个百分点。结果证明结构化运动先验与Transformer表征互补,微小但精调的运动头能显著增强劣质影片分析鲁棒性。相关资源见https://github.com/linty5/DGME-T。

原文摘要 · Abstract (English)

Camera movement classification (CMC) models trained on contemporary, high-quality footage often degrade when applied to archival film, where noise, missing frames, and low contrast obscure motion cues. We bridge this gap by assembling a unified benchmark that consolidates two modern corpora into four canonical classes and restructures the HISTORIAN collection into five balanced categories. Building on this benchmark, we introduce DGME-T, a lightweight extension to the Video Swin Transformer that injects directional grid motion encoding, derived from optical flow, via a learnable and normalised late-fusion layer. DGME-T raises the backbone's top-1 accuracy from 81.78% to 86.14% and its macro F1 from 82.08% to 87.81% on modern clips, while still improving the demanding World-War-II footage from 83.43% to 84.62% accuracy and from 81.72% to 82.63% macro F1. A cross-domain study further shows that an intermediate fine-tuning stage on modern data increases historical performance by more than five percentage points. These results demonstrate that structured motion priors and transformer representations are complementary and that even a small, carefully calibrated motion head can substantially enhance robustness in degraded film analysis. Related resources are available at https://github.com/linty5/DGME-T.

视频分类运动编码历史影像Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。