arXiv:2511.03725cs.CV2025-11NeurIPS被引 3

分离动作与场景概念,让视频动作识别更可解释。

Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition

论文配图:Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
图 1 · 摘自论文原文
  • 将动作分解为运动、物体、场景三类独立概念进行预测
  • 在四个数据集上解释力显著提升,性能保持竞争力
  • 适合需要模型可解释性与调试的开发者

视频动作识别的有效解释应将时间上的运动变化与空间上下文分离。然而,现有基于显著性的方法产生纠缠的解释,难以判断预测依赖运动还是上下文。语言方法虽具结构,却常无法解释运动细节,因其隐含性——直观但难言传。为此,我们提出DANCE框架,通过解耦的动作与上下文概念(运动动态、物体、场景)进行动作预测。运动动态定义为人体姿态序列,物体与场景概念由大语言模型自动提取。基于事前概念瓶颈设计,DANCE强制模型通过这些概念进行决策。在KTH、Penn Action、HAA500和UCF-101四个数据集上的实验表明,DANCE显著提升解释清晰度,同时保持良好性能。用户研究验证其解释性优势。结果还显示DANCE有助于模型调试、编辑与失败分析。

原文摘要 · Abstract (English)

Effective explanations of video action recognition models should disentangle how movements unfold over time from the surrounding spatial context. However, existing methods based on saliency produce entangled explanations, making it unclear whether predictions rely on motion or spatial context. Language-based approaches offer structure but often fail to explain motions due to their tacit nature -- intuitively understood but difficult to verbalize. To address these challenges, we propose Disentangled Action aNd Context concept-based Explainable (DANCE) video action recognition, a framework that predicts actions through disentangled concept types: motion dynamics, objects, and scenes. We define motion dynamics concepts as human pose sequences. We employ a large language model to automatically extract object and scene concepts. Built on an ante-hoc concept bottleneck design, DANCE enforces prediction through these concepts. Experiments on four datasets -- KTH, Penn Action, HAA500, and UCF-101 -- demonstrate that DANCE significantly improves explanation clarity with competitive performance. We validate the superior interpretability of DANCE through a user study. Experimental results also show that DANCE is beneficial for model debugging, editing, and failure analysis.

视频识别可解释性概念分离大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。