零样本骨骼动作预测,让机器人提前识破未知动作意图。
Zero-Shot Skeleton-Based Action Anticipation

- 用互信息对齐视觉特征与语义嵌入,提升未见动作识别能力。
- 在NTU RGB+D数据集上实现高零样本准确率,验证方法有效性。
- 适合需要泛化到新动作的机器人系统研究者参考。
动作预测(AA)旨在从部分观测中识别正在进行的人类或人形机器人动作,使机器人能在动作完成前预判意图。尽管基于骨骼的AA具有效率优势,现有方法假设训练时已见过所有动作类别,限制了其在真实场景中的部署,因新动作不可避免会出现。为此,我们提出新的零样本骨骼动作预测(ZS-SkAA)任务:仅用少量早期骨骼序列识别未见过的动作类别,融合部分观测、时序动态与零样本泛化挑战。为建立基础研究,我们引入:(1) 一个基线模型,包含时空特征提取器与互信息估计最大化模块,通过显式对齐跨模态的局部视觉特征与语义类别嵌入,增强对未见类别的泛化能力;(2) 基于NTU RGB+D数据集的基准评估协议,用于严格评估ZS-SkAA。实验表明,该模型作为强基线,在NTU RGB+D上实现高零样本准确率。本工作确立了ZS-SkAA作为真实系统中应对新动作的关键研究方向。
原文摘要 · Abstract (English)
Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which limits their deployment in real-world scenarios where novel actions inevitably arise. To address this gap, we study the new task of Zero-Shot Skeleton-Based Action Anticipation (ZS-SkAA). This task requires recognizing unseen action classes using only limited early-stage skeleton sequences, combining the challenges of partial observations, temporal dynamics, and zero-shot generalization. To establish foundational research for ZS-SkAA, we introduce:(1) A baseline model comprising a spatio-temporal feature extractor and a mutual information estimation and maximization module. This baseline model explicitly aligns partial visual features with semantic class embeddings across modalities by estimating and maximizing their mutual information, enhancing generalization to unseen classes.(2) A benchmark protocol using the NTU RGB+D dataset, which is adapted for rigorous ZS-SkAA evaluation. Experiments demonstrate the effectiveness of our model as a strong baseline for ZS-SkAA, achieving high zero-shot accuracy on NTU RGB+D. This work establishes ZS-SkAA as a vital research direction for real-world systems requiring generalization to novel actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。