arXiv:2507.16287cs.CV2025-07ICCV被引 3

用大模型拆解动作本质,提升少样本动作识别效果

Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

  • 用大模型将动作标签分解为主体-动作-对象三要素
  • 视频分段捕捉动作时序结构,与文本特征逐级融合
  • 适合需要细粒度动作理解的少样本场景

少样本动作识别(FSAR)旨在仅用少量标注样本对视频中的人类动作进行分类。由于训练数据稀缺,近期研究尝试引入文本等多模态信息。然而,人体姿态、运动动态及物体交互等动作内在细节,在不同阶段变化细微,仅靠动作标签难以充分挖掘。本文提出语言引导的动作解剖框架(LGA),通过大语言模型(LLM)深入解析动作标签的潜在表征。针对文本,提示预训练大模型将标签拆解为包含主体、动作、对象的原子动作序列;针对视频,设计视觉解剖模块将其分割为原子动作阶段,捕捉动作时序结构。通过细粒度融合策略,在原子层级整合图文特征,生成更具泛化性的原型。最后引入视频-视频与视频-文本双重匹配机制,保障少样本分类鲁棒性。实验表明,LGA在多个主流FSAR基准上达到领先性能。

原文摘要 · Abstract (English)

Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, motion dynamics, and the object interactions that occur during different phases, are critical inherent knowledge of actions that cannot be fully exploited by action labels alone. In this work, we propose Language-Guided Action Anatomy (LGA), a novel framework that goes beyond label semantics by leveraging Large Language Models (LLMs) to dissect the essential representational characteristics hidden beneath action labels. Guided by the prior knowledge encoded in LLM, LGA effectively captures rich spatiotemporal cues in few-shot scenarios. Specifically, for text, we prompt an off-the-shelf LLM to anatomize labels into sequences of atomic action descriptions, focusing on the three core elements of action (subject, motion, object). For videos, a Visual Anatomy Module segments actions into atomic video phases to capture the sequential structure of actions. A fine-grained fusion strategy then integrates textual and visual features at the atomic level, resulting in more generalizable prototypes. Finally, we introduce a Multimodal Matching mechanism, comprising both video-video and video-text matching, to ensure robust few-shot classification. Experimental results demonstrate that LGA achieves state-of-the-art performance across multipe FSAR benchmarks.

少样本识别动作解剖多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。