arXiv:2603.06181cs.CV2026-03被引 4

提出运动图灵测试,用动作数据评估机器人是否像人。

Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid Robots

  • 用SMPL-X表示动作,仅凭运动信息判断人类与机器人的区别。
  • 1000段动作数据中,机器人在跳跃等动态动作上仍明显异常。
  • 新基准模型比大语言模型更擅长自动评估动作自然度。

类人机器人在运动生成与控制方面取得显著进展,展现出越来越自然、类人的动作。受图灵测试启发,本文提出运动图灵测试框架,通过仅基于运动学信息,评估人类观察者能否区分类人机器人与人类的动作。为此,我们构建了包含1000个运动序列的HHMotion数据集,覆盖15种动作类别,由11个类人模型和10名人类参与者完成。所有动作均转换为SMPL-X表示,以消除视觉外观影响。我们招募30名标注员对每个动作的人类相似度进行0-5分评分,累计超过500小时标注。数据分析显示,类人机器人动作仍存在明显偏差,尤其在跳跃、拳击和跑步等动态动作中。基于此,我们建立了一个自动预测人类相似度的任务。尽管多模态大语言模型已有进展,但其在动作人类相似度评估上仍表现不足。为此,我们提出一个简单基线模型,并证明其优于多个主流的LLM方法。相关数据集、代码与基准将公开发布,以支持社区后续研究。

原文摘要 · Abstract (English)

Humanoid robots have achieved significant progress in motion generation and control, exhibiting movements that appear increasingly natural and human-like. Inspired by the Turing Test, we propose the Motion Turing Test, a framework that evaluates whether human observers can discriminate between humanoid robot and human poses using only kinematic information. To facilitate this evaluation, we present the Human-Humanoid Motion (HHMotion) dataset, which consists of 1,000 motion sequences spanning 15 action categories, performed by 11 humanoid models and 10 human subjects. All motion sequences are converted into SMPL-X representations to eliminate the influence of visual appearance. We recruited 30 annotators to rate the human-likeness of each pose on a 0-5 scale, resulting in over 500 hours of annotation. Analysis of the collected data reveals that humanoid motions still exhibit noticeable deviations from human movements, particularly in dynamic actions such as jumping, boxing, and running. Building on HHMotion, we formulate a human-likeness evaluation task that aims to automatically predict human-likeness scores from motion data. Despite recent progress in multimodal large language models, we find that they remain inadequate for assessing motion human-likeness. To address this, we propose a simple baseline model and demonstrate that it outperforms several contemporary LLM-based methods. The dataset, code, and benchmark will be publicly released to support future research in the community.

类人机器人动作评估运动理解人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。