评测5个语音驱动手势生成系统在真实对话中的表现。
The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

- 采用解耦评估法分离动作质量和语音对齐度,避免干扰。
- 顶级系统语音对齐率仅32%,远低于真人动作62%的天花板。
- 现有模型无法有效回应对话伙伴,语义表达能力几乎为零。
本文报告了第四届GENEA挑战赛的结果,对五个团队在无缝交互数据集上训练的语音驱动手势生成系统进行了大规模人类评估。延续2023年方法,采用解耦评估机制,分别衡量动作质量与语音对齐性,并通过双人错配实验隔离听觉反应的影响。新增基于地面化手势子集的语义手势生成任务及文本错配评估方法。共开展四项大规模用户研究,收集869名参与者超过23,000次投票。在动作真实度评估中,数据集筛选片段的胜率高达68%-95%;语音对齐方面,动作捕捉片段达到62%的上限,最高提交结果仅为32%,其余接近随机水平(0%);在双人互动评估中,动作捕捉达到65%合适度,但所有系统均未显著优于随机;语义错配评估显示数据集手势具有高度语义表达力(79%正确识别),而系统表现极差,最佳仅达8%合适度。所有投票数据与生成输出将公开于https://genea-workshop.github.io/2026/challenge/,以促进可复现性与后续研究。
原文摘要 · Abstract (English)
This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset's filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment study, the motion-capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input-independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test-takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea-workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。