arXiv:2606.09390cs.CVcs.AI2026-06

用人体姿态识别远距离人机通信意图,兼顾实时性与可靠性。

Real-time body pose non-verbal communication with a consistency-based reliability measure

论文配图:Real-time body pose non-verbal communication with a consistency-based reliability measure
图 1 · 摘自论文原文
  • 基于2D人体姿态识别十类沟通意图,构建新数据集。
  • 在嵌入式GPU上实现实时推理,模型自一致性提升可靠性判断。
  • 适合救援等远程低资源场景的机器人交互系统使用。

身体动作可在无法捕捉面部或语音的远距离环境中传递意图。本文研究仅凭二维人体姿态识别沟通意图的问题。在需实时、低成本、本地部署的人机通信场景(如救援任务)中,身体运动具有高可靠性。然而现有数据集未能专门分离此信号:情感语料库融合面部、语音和文本,而骨架动作识别基准标注的是行为而非传达的信息。为此,我们发布了一个包含十类沟通意图的真实全身姿态帧数据集,并与真实数据(IPC)及合成数据(MotionLCM、VEO3.1、Kimodo)进行对比,覆盖不同难度。针对机器人有限算力,我们测试多种模型(从骨架图分类器到关节运动预测网络),并在NVIDIA Orin Nano嵌入式GPU上报告性能与帧率。最后,发现模型自身自回归自一致性可作为无监督可靠性信号。我们给出理论证明,表明一致步数越多,预测正确的概率越高,并识别出即使自信预测也可能错误的情形,通过行业标准指标验证。

原文摘要 · Abstract (English)

Body movement communicates intent at distances and in conditions where neither the face, nor speech can be captured. We study the recognition of communicative intent from 2D body pose alone. We argue that body motion is a reliable signal especially in scenarios that require real time low-cost on-device person-to-robot communication in long distance environments, such as rescue missions. However, existing resources do not isolate this signal. Affective corpora combine body, face, voice and text, while skeleton action-recognition benchmarks label the action performed rather than the message conveyed. We release a dataset of real frames of full-body pose covering ten communicative intents and we compare it against other real (IPC) and synthetic (MotionLCM, VEO3.1, Kimodo) ones that span a range of difficulty. We target systems that can run on a robot's limited onboard hardware. We benchmark multiple models, from skeleton graph classifiers to joint motion-forecasting networks, and report performance metrics together with frame rate on an embedded GPU (NVIDIA Orin~Nano), since speed matters as much as accuracy in our scenario. Finally, we show that a model's own autoregressive self-consistency works as an unsupervised reliability signal. We give a short proof that bounds the probability that a self-consistent prediction is correct, show that this probability grows with the number of consistent steps, and identify the conditions under which a confident prediction can still be false, benchmarked against industry-standard metrics.

人体姿态人机交互实时系统可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。