arXiv:2604.27508cs.RO2026-04

通过子动作语义提升机器人早期动作识别能力

SASI: Leveraging Sub-Action Semantics for Robust Early Action Recognition in Human-Robot Interaction

论文配图:SASI: Leveraging Sub-Action Semantics for Robust Early Action Recognition in Human-Robot Interaction
图 1 · 摘自论文原文
  • 融合骨架图卷积与子动作语义,实现细粒度动作解析
  • 在BABEL数据集上达到29帧/秒实时处理,早识别准确率显著提升
  • 适合需要快速响应的机器人交互场景,如协作制造、助老服务

理解人类动作对推进人机交互中的行为分析至关重要。在需快速主动反馈的任务中,机器人必须从不完整观测中尽早识别人类动作。子动作提供了所需的语义与层级线索,因为人类动作本质上具有结构,可分解为更小且有意义的单元。然而,传统方法主要关注整体动作,常忽略子动作中蕴含的丰富语义结构,难以支持早期识别。为此,本文提出SASI(子动作语义融合跨模态融合框架),将现有图卷积网络与子动作语义融合,利用基于骨架的分割模型捕捉细粒度子动作语义与整体空间上下文,实现实时运行(29 Hz)。在带有帧级标注的骨架数据集BABEL上的实验表明,该方法优于传统方法,且随着子动作分割质量提升,性能还将进一步提高。特别地,SASI在部分动作序列上表现优异,展现出强大的早期识别能力,对实现主动、无缝的人机交互至关重要。代码已公开于https://anonymous.4open.science/r/SASI。

原文摘要 · Abstract (English)

Understanding human actions is critical for advancing behavior analysis in human-robot interaction. Particularly in tasks that demand quick and proactive feedback, robots must recognize human actions as early as possible from incomplete observations. \textit{Sub-actions} offer the semantic and hierarchical cues needed for this, since human actions are inherently structured and can be decomposed into smaller, meaningful units. However, conventional approaches focus primarily on holistic actions and often overlook the rich semantic structure embedded in sub-actions, making them poorly suited for early recognition. To address this gap, we introduce SASI (Sub-Action Semantics Integrated cross-modal fusion), a novel framework that integrates existing graph convolution networks to fuse spatiotemporal features with sub-action semantics. SASI exploits a segmentation model with a traditional skeleton-based graph convolution network, capturing both fine-grained sub-action semantics and overall spatial context, while operating in real-time at 29 Hz. Experiments on BABEL, a skeleton-based dataset with frame-level annotations, demonstrate that our method improves recognition accuracy over conventional approaches, with additional gains expected as the quality of sub-action segmentation improves. Notably, SASI also achieves superior performance in understanding partial action sequences, revealing its capability for early recognition, which is essential for proactive and seamless Human-Robot Interaction (HRI). Code is available at https://anonymous.4open.science/r/SASI .

动作识别人机交互子动作实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。