arXiv:2511.09388cs.CV2025-11中稿 · CVPR被引 3

用邻域语义与开放流匹配,提升零样本骨骼动作识别鲁棒性。

Learning by Neighbor-Aware Semantics, Deciding by Open-form Flows: Towards Robust Zero-Shot Skeleton Action Recognition

  • 引入邻域上下文增强文本语义,实现点到区域的动态对齐。
  • 在仅用10%训练数据下,三个数据集上均显著优于现有方法。
  • 适合关注零样本动作识别与跨模态对齐的研究者。

识别未见骨骼动作类别仍具挑战性,因缺乏对应骨骼先验。现有方法多采用“对齐-分类”范式,但存在两点根本问题:(i)语义不完善导致点对点对齐脆弱;(ii)分类器受限于静态决策边界和粗粒度锚点。为此,本文提出新方法 exttt{Flora},基于灵活的邻域感知语义调谐与开放形式分布感知流分类器。具体而言,通过融合邻近类别的上下文线索构建方向感知的区域语义,并引入跨模态几何一致性目标,实现稳定可靠的点到区域对齐。同时,采用无噪流匹配弥合语义与骨骼嵌入间的模态分布差异,条件无关对比正则化增强判别力,通过令牌级速度预测实现细粒度决策边界。在三个基准数据集上的实验验证了方法有效性,尤其在仅使用10%已见数据训练时表现优异。代码已公开于 https://github.com/cseeyangchen/Flora。

原文摘要 · Abstract (English)

Recognizing unseen skeleton action categories remains highly challenging due to the absence of corresponding skeletal priors. Existing approaches generally follow an ``align-then-classify'' paradigm but face two fundamental issues, \textit{i.e.}, (i) fragile point-to-point alignment arising from imperfect semantics, and (ii) rigid classifiers restricted by static decision boundaries and coarse-grained anchors. To address these issues, we propose a novel method for zero-shot skeleton action recognition, termed \texttt{\textbf{Flora}}, which builds upon \textbf{F}lexib\textbf{L}e neighb\textbf{O}r-aware semantic attunement and open-form dist\textbf{R}ibution-aware flow cl\textbf{A}ssifier. Specifically, we flexibly attune textual semantics by incorporating neighboring inter-class contextual cues to form direction-aware regional semantics, coupled with a cross-modal geometric consistency objective that ensures stable and robust point-to-region alignment. Furthermore, we employ noise-free flow matching to bridge the modality distribution gap between semantic and skeleton latent embeddings, while a condition-free contrastive regularization enhances discriminability, leading to a distribution-aware classifier with fine-grained decision boundaries achieved through token-level velocity predictions. Extensive experiments on three benchmark datasets validate the effectiveness of our method, showing particularly impressive performance even when trained with only 10% of the seen data. Code is available at https://github.com/cseeyangchen/Flora.

零样本识别骨骼动作跨模态对齐分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。