构建首个互动舞蹈动作数据集,评估机器人动作是否符合社交语境与对方水平。
CoMPAS3D: A Dataset and Benchmark for Interactive Motion
- 用真人舞者实录3小时即兴双人萨尔萨,标注2800+动作片段
- 提出可衡量动作可懂性与水平匹配度的新评估指标
- 适合研究人机协作、具身智能与动作生成的团队使用
社交互动类人机器人需通过身体实时适应搭档的动作、意图与能力。这要求模型不仅理解动作本身,还需理解其在共享社交语境中的意义。然而现有互动动作生成评估框架无法衡量生成动作是否在共享动作词汇中可懂,也无法判断是否适配搭档水平。这一空白源于两方面:现有框架依赖FID、节拍对齐等运动学指标,无法捕捉上述属性;且现有数据集缺乏动作标注与水平差异。萨尔萨舞具备即兴、双人、有动作词汇与评判标准(时间感、音乐性、技巧、难度、配合、创意)等特性,是理想评估场景。本文提出CoMPAS3D,包含18名舞者(初、中、高阶)共3小时即兴双人萨尔萨动作捕捉数据,超过2,800个专家标注段落,涵盖动作类型、错误与风格元素。构建涵盖运动学质量、两个客观指标(动作可懂性、水平匹配度)和六个竞赛式主观维度的评估框架。设立三项基准:动作分类(类比转录)、水平估计(流利度评估)、跟舞生成(对话响应)。微调视觉-语言模型在真实动作序列上表现良好。应用于Duolando与InterGen时,新指标揭示了运动学指标忽略的缺陷。人工评估确认生成动作与真实动作间的差距。CoMPAS3D数据、标注、基准代码与基线结果均已公开。
原文摘要 · Abstract (English)
Socially interactive humanoid robots must engage with humans through their bodies, adapting in real time to a partner's movement, intent, and abilities. This requires models that understand not just how bodies move, but what movement means in a shared social context. Yet evaluation frameworks for interactive motion generation do not measure whether generated follower motion is legible within a shared movement vocabulary, nor whether it is appropriate to the partner's proficiency level. This gap has two causes: existing frameworks rely on kinematic metrics such as FID and beat alignment that cannot measure either property, and existing datasets lack the move annotations and proficiency variation needed. Salsa is well-suited as an evaluation domain: improvised, dyadic, and governed by a move vocabulary and judging criteria covering timing, musicality, technique, difficulty, partnering, and originality. We present CoMPAS3D, a motion capture dataset of improvised partner salsa paired with an evaluation framework covering kinematic quality, two objective metrics (move legibility and proficiency appropriateness), and six competition-based subjective dimensions. The dataset includes 3 hours of improvisation by 18 dancers spanning beginner, intermediate, and professional levels, with over 2,800 expert-annotated segments covering move types, errors, and stylistic elements. We define three benchmarks: move classification (analogous to transcription), proficiency estimation (fluency assessment), and follower generation (dialogue response). Fine-tuned vision-language models perform strongly on objective metrics applied to ground-truth motion sequences. Applied to Duolando and InterGen, the metrics reveal failures that kinematic metrics miss. Human evaluations confirm the gap between generated and ground-truth motion. CoMPAS3D, annotations, benchmark code, and baseline results are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。