arXiv:2606.06853cs.CVcs.AI2026-06中稿 · CVPR

用视频扩散模型增强视觉语言模型的运动理解能力

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

论文配图:MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models
图 1 · 摘自论文原文
  • 从视频扩散模型中提取运动先验,通过注意力对齐提升视觉语言模型
  • 在两个运动级视频理解基准上显著提升性能,尤其在运动相关指标上
  • 无需额外参数或修改架构,适合快速部署到现有模型

当前视觉语言模型在事件或故事级理解上表现优异,但在捕捉细粒度运动细节方面仍显不足,主要因其侧重高层静态语义结构和宏观事件逻辑。相比之下,视频扩散模型(VDMs)能有效建模动态运动模式,得益于大规模视频数据和时序生成需求。本文提出MotionEnhancer,一种利用强大视频扩散模型中的运动先验作为辅助监督的方法,通过注意力对齐增强视觉语言模型(VLM)的运动理解能力。该方法包含两个无参模块:运动敏感头选择(MHS)和运动显著文本标记识别(MTTI),可直接在计算层面提取并优化与运动相关的注意力。MotionEnhancer无需新增训练参数、架构修改或工具调用,是一种可扩展的运动理解解决方案。大量实验表明,其在两个运动级视频理解基准上均优于现有先进VLM,在运动相关指标上提升显著。

原文摘要 · Abstract (English)

The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-level understanding, their ability to capture fine-grained motion details remains limited, primarily due to their focus on high-level static semantic structures and macro-event logic. In contrast, Video Diffusion Models (VDMs) are adept at modeling dynamic motion patterns, benefiting from large-scale video data and the intrinsic requirement of temporal generation. In this paper, we introduce MotionEnhancer, a novel approach that leverages motion priors distilled from a powerful video diffusion model as auxiliary supervision to enhance the motion understanding capability of a VLM via attention alignment. MotionEnhancer comprises two simple parameter-free modules, Motion-sensitive Head Selection (MHS) and Motion-salient Text Token Identification (MTTI), to directly extract and optimize motion-related attentions from the VDM in a computation-only manner. MotionEnhancer provides a scalable solution for motion understanding without additional training parameters, modifications to existing architectures, or tool calling. Extensive experiments demonstrate that MotionEnhancer can achieve consistent improvements over state-of-the-art VLMs on two motion-level video understanding benchmarks, especially on motion-related metrics.

视觉语言模型视频扩散模型运动理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。