arXiv:2605.07140cs.CVcs.AI2026-05中稿 · Proceedings of the…被引 1

用逻辑规则解释动作识别,让模型推理过程可读可懂。

Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition

论文配图:Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
图 1 · 摘自论文原文
  • 将动作识别转为基于运动单元的逻辑推理
  • 在NTU RGB+D和NW-UCLA上达到先进性能
  • 结合大模型描述构建可解释的概念空间

基于骨架的人体动作识别虽有良好表现,但多数模型缺乏可解释性。本文提出一种神经符号框架,将动作识别重构为基于运动基元的首阶逻辑推理。通过时空骨架编码器提取隐式运动表示,并经由显式分离姿态与动态抽象的时空概念解码器映射为可解释的概念谓词。这些谓词通过可微分首阶逻辑层组合,学习可读的语义规则。为建立语义结构,将骨架表示对齐大语言模型生成的原子运动基元描述,形成感知与推理共享的概念空间。在NTU RGB+D 60/120和NW-UCLA数据集上的实验表明,该方法在保持竞争性识别性能的同时,提供基于逻辑结构的明确解释。结果表明神经符号推理是实现可解释时空动作理解的有效范式。

原文摘要 · Abstract (English)

Skeleton-based human activity recognition has achieved strong empirical performance, yet most existing models remain black boxes and difficult to interpret. In this work, we introduce a neurosymbolic formulation of skeleton-based HAR that reframes action recognition as concept-driven first-order logical reasoning over motion primitives. Our framework bridges representation learning and symbolic inference by grounding first-order logic predicates in learnable spatial and temporal motion concepts. Specifically, we employ a standard spatio-temporal skeleton encoder to extract latent motion representations, which are then mapped to interpretable concept predicates via a spatio-temporal concept decoder that explicitly separates pose-centric and dynamics-centric abstractions. These concept predicates are composed through differentiable first-order logic layers, enabling the model to learn human-readable logical rules that govern action semantics. To impose semantic structure on the learned concepts, we align skeleton representations with LLM-derived descriptions of atomic motion primitives, establishing a shared conceptual space for perception and reasoning. Extensive experiments on NTU RGB+D 60/120 and NW-UCLA demonstrate that our approach achieves competitive recognition performance while providing explicit, interpretable explanations grounded in logical structure. Our results highlight neurosymbolic reasoning as an effective paradigm for interpretable spatio-temporal action understanding. Code: https://github.com/Mr-TalhaIlyas/REASON

动作识别可解释性逻辑推理神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。