用视觉提示提升多模态模型对视频细微运动的理解能力。
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs

- 引入物体中心光斑和运动模糊作为零样本视觉提示。
- 在40K视频片段上实现领先开源性能,媲美商用模型。
- 适合研究视频细粒度运动建模或数据构建的学者使用。
尽管多模态大语言模型(MLLMs)取得了进展,其在细粒度视频运动理解方面仍存在严重局限。它们常缺乏帧间差异分析,往往忽略或平均微弱视觉线索。此外,虽然视觉提示在静态图像中展现潜力,但其在视频时序复杂性中的应用,尤其是细粒度运动理解方面仍几乎未被探索。本研究探讨是否可激发内在能力,增强MLLMs的运动感知,并实现区分物体与相机运动的视觉特征。为此,我们提出MotionSight——一种新颖的零样本方法,首次引入物体中心视觉光斑和运动模糊作为视觉提示,无需训练即可显著提升细粒度运动理解。为支持该研究,我们构建了首个大规模细粒度视频运动理解数据集MotionVid-QA,包含约40K视频片段和约87K个问答对,具有分层标注(SFT与偏好数据)。实验表明,MotionSight达到开源模型最优水平,且与商业模型竞争力相当。相关代码与标注将公开。
原文摘要 · Abstract (English)
Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potential in static images, its application to video's temporal complexities, particularly for fine-grained motion understanding, remains largely unexplored. We investigate whether inherent capability can be unlocked and boost MLLMs' motion perception and enable distinct visual signatures tailored to decouple object and camera motion cues. In this study, we introduce MotionSight, a novel zero-shot method pioneering object-centric visual spotlight and motion blur as visual prompts to effectively improve fine-grained motion understanding without training. To convert this into valuable data assets, we curated MotionVid-QA, the first large-scale dataset for fine-grained video motion understanding, with hierarchical annotations including SFT and preference data, Θ(40K) video clips and Θ(87K) QAs. Experiments show MotionSight achieves state-of-the-art open-source performance and competitiveness with commercial models. In particular, for fine-grained motion understanding we present a novel zero-shot technique and a large-scale, high-quality dataset. All the code and annotations will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。