arXiv:2503.17690cs.CV2025-03CVPR被引 15

用大模型+文本提示实现可泛化的重复动作计数。

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

  • 用大语言模型结合周期性文本提示进行计数
  • 在新动作和域外数据上表现显著优于传统方法
  • 适合需要强泛化能力的动作分析场景

重复动作计数旨在统计视频中的周期性运动,对健身监测等应用至关重要。现有方法多依赖表示能力有限的回归网络,且在狭窄训练集上采用监督学习,易过拟合,泛化能力差。为此,我们提出 CountLLM,首个基于大语言模型(LLM)的框架,输入视频与周期性文本提示,输出计数结果。CountLLM 利用显式文本指令中的丰富线索及预训练 LLM 的强大表征能力。为有效引导模型,我们设计基于周期性的结构化提示模板,描述周期属性并统一答案格式。此外,提出渐进式多模态训练范式,增强模型对周期性的感知。在多个主流基准上的实证评估表明,CountLLM 在性能和泛化性上均更优,尤其在处理与训练数据差异显著的新动作和域外动作时表现突出,为重复动作计数提供了新路径。

原文摘要 · Abstract (English)

Repetitive action counting, which aims to count periodic movements in a video, is valuable for video analysis applications such as fitness monitoring. However, existing methods largely rely on regression networks with limited representational capacity, which hampers their ability to accurately capture variable periodic patterns. Additionally, their supervised learning on narrow, limited training sets leads to overfitting and restricts their ability to generalize across diverse scenarios. To address these challenges, we propose CountLLM, the first large language model (LLM)-based framework that takes video data and periodic text prompts as inputs and outputs the desired counting value. CountLLM leverages the rich clues from explicit textual instructions and the powerful representational capabilities of pre-trained LLMs for repetitive action counting. To effectively guide CountLLM, we develop a periodicity-based structured template for instructions that describes the properties of periodicity and implements a standardized answer format to ensure consistency. Additionally, we propose a progressive multimodal training paradigm to enhance the periodicity-awareness of the LLM. Empirical evaluations on widely recognized benchmarks demonstrate CountLLM's superior performance and generalization, particularly in handling novel and out-of-domain actions that deviate significantly from the training data, offering a promising avenue for repetitive action counting.

动作计数大模型视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。