arXiv:2608.25495cs.CVcs.AI2026-08

用人体姿态锚定光流,提前预测人类动作,适合机器人实时协作。

Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming

论文配图:Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming
图 1 · 摘自论文原文
  • 以人体姿态为参考,聚焦关节周围局部运动,构建结构化光流表示。
  • 在早期观测时(如10%动作序列)仍保持高准确率,优于传统方法。
  • 无需全图光流计算,适合资源受限的实时机器人系统使用。

人机交互(HRI)要求机器人在人类动作初期就理解其意图,以便安全、高效、自然地响应。现有方法或依赖稀疏骨骼信息(缺乏细粒度运动线索),或使用密集光流(计算开销大,难用于低延迟系统)。本文提出PoseOFF——一种以人体姿态为中心的光流表示方法,通过在关键关节附近捕捉局部运动信息,实现与人体运动学对齐的结构化运动表征。该方法在多个基准数据集和骨干网络上验证,均在早期观察比例下(如仅观测动作前10%)取得一致性能提升,达到或超越现有模型表现。更重要的是,其无需全帧光流处理,适用于实时与资源受限场景。结果表明,基于姿态的运动表示能显著提升机器人系统对人类动作的早期感知能力,增强交互中的响应性与前瞻性。

原文摘要 · Abstract (English)

Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human action recognition rely either on sparse skeletal representations, which lack fine-grained motion cues, or dense optical flow, which can be computationally expensive for low-latency perception pipelines. In this paper, we propose PoseOFF, a pose-anchored optical flow representation that captures local motion information around human joints to support earlier human intent understanding. By conditioning motion feature extraction on human pose, PoseOFF encodes localised motion dynamics at semantically meaningful body locations, forming a structured motion representation that is explicitly aligned with human kinematics. We evaluate PoseOFF across multiple benchmark datasets and backbone architectures for action anticipation, demonstrating consistent improvements in recognition accuracy, particularly at early observation ratios. Our results show that PoseOFF enables models to achieve comparable or improved performance while observing less of the action sequence, highlighting its effectiveness for early prediction. Importantly, these gains are achieved without requiring full-frame motion processing, making the approach practical for real-time and resource-constrained settings. These findings suggest that pose-centred motion representations such as PoseOFF can enhance the ability of interactive robot systems to infer human actions earlier, supporting more responsive and anticipatory behaviour in human-robot interaction scenarios.

动作预测人机交互光流实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。