arXiv:2504.13140cs.CV2025-04CVPR被引 2

用人体姿态序列解释动作识别,让模型决策更透明可懂。

PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition

论文配图:PCBEAR: Pose Concept Bottleneck for Explainable Action Recognition
图 1 · 摘自论文原文
  • 以人体骨骼姿态作为结构化概念,捕捉运动动态
  • 在KTH、Penn-Action等数据集上实现高准确率与可解释性
  • 无需人工标注,自动发现有意义的运动模式

人类动作识别(HAR)在深度学习模型下取得显著进展,但其决策过程因黑箱特性而难以理解。确保可解释性对需透明与问责的真实场景至关重要。现有视频XAI方法多依赖像素级特征或静态文本概念,难以捕捉动作理解所需的动作动态与时间依赖。为此,我们提出姿态概念瓶颈框架PCBEAR,引入人体姿态序列作为运动感知、结构化的概念,用于视频动作识别。不同于基于像素或静态描述的方法,PCBEAR聚焦人体骨架姿态,仅关注身体运动,提供对动作动态的鲁棒且可解释的说明。我们定义两类姿态概念:单帧空间构型的静态姿态概念,以及跨帧运动模式的动态姿态概念。通过聚类视频姿态序列构建这些概念,实现无需人工标注的自动发现。我们在KTH、Penn-Action和HAA500数据集上验证了PCBEAR,结果表明其在保持高分类性能的同时,提供可解释的运动驱动推理。该方法兼具强预测能力与人类可理解的推理洞察,支持测试时干预以调试和改进模型行为。

原文摘要 · Abstract (English)

Human action recognition (HAR) has achieved impressive results with deep learning models, but their decision-making process remains opaque due to their black-box nature. Ensuring interpretability is crucial, especially for real-world applications requiring transparency and accountability. Existing video XAI methods primarily rely on feature attribution or static textual concepts, both of which struggle to capture motion dynamics and temporal dependencies essential for action understanding. To address these challenges, we propose Pose Concept Bottleneck for Explainable Action Recognition (PCBEAR), a novel concept bottleneck framework that introduces human pose sequences as motion-aware, structured concepts for video action recognition. Unlike methods based on pixel-level features or static textual descriptions, PCBEAR leverages human skeleton poses, which focus solely on body movements, providing robust and interpretable explanations of motion dynamics. We define two types of pose-based concepts: static pose concepts for spatial configurations at individual frames, and dynamic pose concepts for motion patterns across multiple frames. To construct these concepts, PCBEAR applies clustering to video pose sequences, allowing for automatic discovery of meaningful concepts without manual annotation. We validate PCBEAR on KTH, Penn-Action, and HAA500, showing that it achieves high classification performance while offering interpretable, motion-driven explanations. Our method provides both strong predictive performance and human-understandable insights into the model's reasoning process, enabling test-time interventions for debugging and improving model behavior.

动作识别可解释性姿态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。