arXiv:2606.03837cs.CV2026-06

研究视频任务适配中时间上下文如何分配更有效

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

  • 对比不同模型组件中时间信息的分布策略
  • 在小数据场景下,时间上下文分配影响适配效果
  • 为参数高效视频建模提供新思路,适合资源受限场景

参数高效微调(PEFT)和探测方法可通过极少可训练参数实现基础模型的适配,适用于标注与计算成本高的视频理解任务。然而,现有视频PEFT主要针对图像预训练模型,而标准PEFT方法同样可用于视频表示。这些设置很少被比较,且均将时序推理局限于模型单一组件,未明确时间上下文应如何分布在骨干网络、PEFT模块和探测器之间。本文系统研究了视频理解中的模型适配策略,评估了在外观关注、运动关注及空间密集场景下的方法表现,尤其聚焦于数据有限场景——此时参数效率最为关键。结果揭示了不同场景下PEFT与探测的有效性差异,并证明时间上下文分配对高效视频适配至关重要。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it attractive for video understanding where annotation and computation are expensive. However, video PEFT has focused on adapting image-pretrained models, while standard PEFT methods can also be applied to video representations. These settings are rarely compared and both confine temporal reasoning to a single component of the model, leaving open how temporal context should be distributed across backbone, PEFT and probe. In this work we provide a systematic study of model adaptation strategies for video understanding. We evaluate methods across appearance-focused, motion-focused and spatially dense settings, with a particular focus on scenarios with limited data where parameter-efficiency is most beneficial. Our results provide new insights into PEFT and probing across settings and demonstrate the importance of temporal context allocation for effective video adaptation

视频理解参数高效时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。