arXiv:2506.16701cs.CV2025-06

用语言先验增强视频动作识别,提升复杂场景理解能力。

Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

  • 引入语言驱动的常识先验,生成场景上下文描述
  • 在Action Genome和Charades数据集上显著提升识别精度
  • 适合需要理解复杂交互的视频分析任务

近期视频动作识别方法通过将大规模预训练的语言-图像模型迁移到视频领域,取得了优异性能。然而,语言模型中蕴含丰富的常识先验——人类理解物体、人-物交互及活动时依赖的场景上下文——尚未被充分挖掘。本文提出一个融合语言驱动常识先验的框架,旨在从单目视角、常被遮挡的杂乱视频序列中识别动作。方法包括:(1) 视频上下文摘要模块,生成候选物体、活动及物体与活动间的交互;(2) 描述生成模块,基于上下文推理后续活动并生成场景描述,利用辅助提示与常识推理;(3) 多模态动作识别头,融合视觉与文本线索进行动作识别。在具有挑战性的Action Genome和Charades数据集上验证了方法的有效性。

原文摘要 · Abstract (English)

Recent video action recognition methods have shown excellent performance by adapting large-scale pre-trained language-image models to the video domain. However, language models contain rich common sense priors - the scene contexts that humans use to constitute an understanding of objects, human-object interactions, and activities - that have not been fully exploited. In this paper, we introduce a framework incorporating language-driven common sense priors to identify cluttered video action sequences from monocular views that are often heavily occluded. We propose: (1) A video context summary component that generates candidate objects, activities, and the interactions between objects and activities; (2) A description generation module that describes the current scene given the context and infers subsequent activities, through auxiliary prompts and common sense reasoning; (3) A multi-modal activity recognition head that combines visual and textual cues to recognize video actions. We demonstrate the effectiveness of our approach on the challenging Action Genome and Charades datasets.

视频动作识别常识推理多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。