arXiv:2506.01119cs.CV2025-06被引 1

用光流提升视频理解的时序建模效率与可解释性

MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows

  • 结合光流与空间嵌入,构建以时间为中心的视频编码架构
  • 在临床、医疗和动作识别数据集上达到顶尖性能
  • 利用预训练模型降低计算成本,适合医疗等敏感场景

许多以运动为核心的视频分析任务,如原子动作识别、自闭症患者异常运动行为检测,或实时磁共振下言语发音运动分析,需要高效且可解释的时序建模。捕捉时序动态是视频分析的核心挑战,通常需大量计算资源和细粒度标注,而这些标注并不广泛可用。本文提出MOOSE(Motion Flow Over Spatial Space),一种新颖的以时间为中心的视频编码器,通过将光流与空间嵌入显式结合,高效建模时序信息,灵感来自人类对运动的感知。与以往方法不同,MOOSE利用广泛可用的预训练视觉和光流编码器,无需从零训练视频模型,显著降低计算复杂度,同时增强时序可解释性。主要贡献包括:(1) 提出一种计算高效的时序中心视频理解架构;(2) 在建模时序动态方面实现更好可解释性;(3) 在涵盖临床、医疗及标准动作识别的多个基准上达到当前最优性能,验证了方法的广泛适用性与有效性。

原文摘要 · Abstract (English)

Many motion-centric video analysis tasks, such as atomic actions, detecting atypical motor behavior in individuals with autism, or analyzing articulatory motion in real-time MRI of human speech, require efficient and interpretable temporal modeling. Capturing temporal dynamics is a central challenge in video analysis, often requiring significant computational resources and fine-grained annotations that are not widely available. This paper presents MOOSE (Motion Flow Over Spatial Space), a novel temporally-centric video encoder explicitly integrating optical flow with spatial embeddings to model temporal information efficiently, inspired by human perception of motion. Unlike prior models, MOOSE takes advantage of rich, widely available pre-trained visual and optical flow encoders instead of training video models from scratch. This significantly reduces computational complexity while enhancing temporal interpretability. Our primary contributions includes (1) proposing a computationally efficient temporally-centric architecture for video understanding (2) demonstrating enhanced interpretability in modeling temporal dynamics; and (3) achieving state-of-the-art performance on diverse benchmarks, including clinical, medical, and standard action recognition datasets, confirming the broad applicability and effectiveness of our approach.

视频理解光流时序建模医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。