arXiv:2510.08480cs.CV2025-10被引 30

用工具增强视频动作识别,让模型更懂细微差别。

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

  • 把动作拆成小片段,精准匹配
  • 引入外部工具提升跨模态推理准确率
  • 适合需要细粒度动作理解的场景

多模态大语言模型在视觉与文本推理间展现出巨大潜力,但其依赖文本先验的特性常限制其在开放词汇场景下区分语义相近动作的能力。为此,我们提出 Video-STAR 框架,通过上下文子动作分解与工具增强的强化学习,实现开放词汇动作识别(OVAR)。不同于将动作视为整体的传统方法,本方案创新性地将动作分解为可区分的子动作,实现细粒度匹配,并动态调用领域专用工具进行跨模态交织,从而增强类别特异性推理能力并减少跨模态幻觉。此外,通过设计分层奖励机制,在工具使用效率、子动作相关性与推理结构一致性之间取得平衡,使模型能自主利用外部工具优先关注子动作模式,无需显式监督,实现从文本中心推理向视觉基础推理的转变。在 HMDB-51、UCF-101、SSv2、Kinetics-400 与 Kinetics-600 等数据集上的大量实验表明,该方法在区分细粒度动作和处理跨模态幻觉方面均优于现有方法,验证了其出色的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text-centric priors often limits their ability to disentangle semantically similar actions in open-vocabulary scenarios. To address this, we propose Video-STAR, a framework that harmonizes contextual sub-motion decomposition with tool-augmented reinforcement learning for open-vocabulary action recognition (OVAR). Unlike prior methods that treat actions as monolithic entities, our approach innovatively decomposes actions into discriminative sub-motions for fine-grained matching while dynamically invoking domain-specific tools for cross-modal interleaving, thereby enabling category-specific reasoning capacity and reducing cross-modal hallucination. Moreover, by designing a hierarchical reward that balances tool-usage efficiency, sub-motion relevance, and structural coherence in reasoning, our method autonomously leverages external tools to prioritize sub-motion patterns without explicit supervision, transmitting from text-centric reasoning to visually grounded inference. Extensive evaluations on HMDB-51, UCF-101, SSv2, Kinetics-400, and Kinetics-600 datasets demonstrate our state-of-the-art performance, outperforming existing methods in distinguishing fine-grained actions and handling cross-modal hallucination, validating our excellent robustness and generalization.

动作识别多模态强化学习细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。