解决视觉受限下骨骼动作识别的缺失关节问题,提升实际场景中的识别鲁棒性。
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach

- 构建带可学习虚拟边的超图,捕捉关节间的高阶依赖关系。
- 引入可见性先验自适应过滤遮挡信息,性能在严重遮挡下提升68.8%。
- 适用于穿戴设备、边缘机器人等真实受限视野场景。
基于骨骼的动作识别在利用关节点坐标及其拓扑连接方面取得显著进展,但现有方法普遍假设输入为完整且清晰的骨骼数据。在实际部署中,如第一人称视觉、拥挤监控、可穿戴设备或边缘机器人场景下,有限视场(FoV)常导致大量关节点不可见,引发性能严重下降,而现有模型对此类问题应对能力不足。为填补这一关键但未被充分研究的空白,我们提出PartialVisGraph,一种专为受限视场下鲁棒骨骼动作识别设计的超图框架。首先,通过引入可学习的虚拟超边构建高度表达的超图,形成软关联矩阵,捕捉超越传统成对图的灵活高阶依赖。随后,提出单头自适应变压器(Single-Head Sample-Adaptive Transformer),在超边上自适应聚合关节点特征,并显式融入可见性先验,选择性抑制遮挡或出视野关节点的信息传播,防止其污染可靠特征传递。我们进一步建立了严格的评估协议,使用真实视场模拟基准在NTU RGB+D 60和120上进行测试。大量实验表明,PartialVisGraph在部分可见条件下持续达到最优准确率,相较于近期强基线,在严重视场限制子集上最高提升达68.8%,同时在全可见设置下仍保持优势。该方法为无约束环境下可部署的骨骼动作理解提供了原则性且实用的路径。
原文摘要 · Abstract (English)
Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In real-world deployments, such as egocentric vision, crowded surveillance, wearable devices, or edge robotics, limited field-of-view (FoV) frequently causes substantial joint visibility dropout, leading to severe performance degradation that existing models are largely unprepared to handle. To bridge this critical yet underexplored gap, we introduce PartialVisGraph, a novel hypergraph framework tailored for robust skeleton action recognition under constrained FoV. We first construct highly expressive hypergraphs by introducing learnable virtual hyperedges that form a soft incidence matrix, capturing flexible high-order dependencies beyond conventional pairwise graphs. We then propose the Single-Head Sample-Adaptive Transformer, which adaptively aggregates joint features onto hyperedges while explicitly incorporating a visibility prior. This prior selectively gates information flow, preventing occluded or out-of-view joints from corrupting reliable feature propagation. We further establish rigorous evaluation protocols with realistic FoV simulation benchmarks on NTU RGB+D 60 and 120. Extensive experiments demonstrate that PartialVisGraph consistently achieves state-of-the-art accuracy under partial visibility, with gains of up to 68.8\% on subsets with severe FoV restrictions compared to recent strong baselines, while remaining superior on full-visibility settings. Our approach offers a principled and practical pathway toward deployable skeleton-based action understanding in unconstrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。