arXiv:2608.13555cs.ROcs.AI2026-08中稿 · ECCV

构建首个兼顾人感知与可扩展性的拟人运动追踪评测基准

HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

论文配图:HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
图 1 · 摘自论文原文
  • 基于真实表演者153小时动作数据,按四类运动细分并标注细节
  • 提出HumanScore评分体系,通过24000条动作对训练,更贴近人类偏好
  • 能识别传统方法忽略的足部滑移、落地时机错误等物理异常

拟人运动追踪在远程操控与全身模仿中至关重要,但现有评估方式常与人类视觉感知不符。传统基于帧级姿态差异的运动学误差指标会忽略关键物理问题,如支撑不稳、接触错误(如足部滑移、落地时机不准)。现有测试集规模小且缺乏多样性,难以检验复杂接触行为和长时序动作。为此,我们提出HumanTracker基准,包含约153小时来自多位专业表演者的光学动捕数据,分为四类运动家族并配有文本标签,支持细粒度诊断。进一步提出HumanScore,一种基于12,000组动作对(共24,000条动作)训练的偏好对齐指标。在主流追踪器上验证表明,HumanScore更能预测人类偏好,并揭示运动学指标常遗漏的接触与稳定性失效问题。

原文摘要 · Abstract (English)

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.

动作追踪评测基准感知对齐运动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。