arXiv:2607.26097cs.CV2026-07

用细粒度语义知识拆解动作,提升复杂场景下动作识别精度

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

论文配图:Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
图 1 · 摘自论文原文
  • 通过大模型将动作标签分解为原子动作,提供时空语义
  • 引入知识注入与解耦模块,使动作特征更清晰可分
  • 适用于多标签动作识别,可无缝融入其他方法

复杂场景中的动作识别常涉及多个并发的细粒度动作,现有方法多依赖整体表征,难以捕捉细微交互与语义。尽管提示学习方法引入了解耦机制,但缺乏显式语义引导;仅依赖视觉或结构线索的方法仍较粗略。本文提出知识引导的原子动作解耦方法(KDA),利用大语言模型(LLM)将动作标签分解为原子动作,赋予明确的时空语义。知识注入模块(KIM)将原子动作知识融合至视频特征;基于此,知识解耦模块(KDM)进一步解耦知识成分,生成更精准的语义引导。引入知识解耦损失(KD Loss)促进组件间清晰解耦。大量实验表明,KDA显著提升特征区分性,在多标签动作识别基准上达到当前最优性能。KIM与KDM可灵活嵌入其他方法,具备强通用性。

原文摘要 · Abstract (English)

Action recognition in complex scenes often involves multiple concurrent fine-grained actions, making it challenging to model internal action structures. Most existing methods rely on holistic representations, which are insufficient for capturing subtle interactions and fine-grained semantics. While recent prompt-based approaches introduce disentanglement, they lack explicit semantic guidance, and methods based solely on visual or structured cues remain coarse-grained. In this paper, we propose Knowledge-guided Disentanglement with Atomic Actions (KDA), which leverages fine-grained semantic knowledge to enhance action representations and enable more precise disentanglement. Specifically, we use Large Language Models (LLMs) to decompose action labels into atomic actions, providing explicit spatial-temporal semantics. A Knowledge Injection Module (KIM) first integrates atomic action knowledge into video features. Based on this enhanced representation, a Knowledge Disentanglement Module (KDM) further disentangles atomic action knowledge to produce more precise semantic guidance for action disentanglement. A Knowledge Disentanglement Loss (KD Loss) is introduced to encourage clearer disentanglement of knowledge components within KDM. Extensive experiments demonstrate that KDA improves feature discriminability and achieves state-of-the-art performance on multi-label action recognition benchmarks. Moreover, KIM and KDM can be readily integrated into other methods, demonstrating strong generality.

动作识别知识引导解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。