建立标准化评估流程,揭示手势生成模型真实性能差距。
Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
- 设计可复现的人类评估协议,统一测试标准。
- 六款模型在动作真实感上表现趋同,近年模型无明显优势。
- 首次公开超750个视频与1.6万份偏好数据,支持后续研究。
我们回顾了自动语音驱动3D手势生成中的人类评估实践,发现缺乏标准化且常采用有缺陷的实验设计,导致无法准确比较不同方法或界定当前技术状态。为解决评估设计常见问题并统一未来手势生成研究中的用户研究,我们针对广泛使用的BEAT2动作捕捉数据集提出详细的人类评估协议。基于该协议,我们开展了大规模众包评估,对六款由原作者训练的近期手势生成模型进行评价,涵盖动作真实感和语音-手势对齐两个关键维度。结果表明:1)在BEAT2数据集上,动作真实感已趋于饱和,旧模型与新模型表现相当;2)此前宣称的高语音-手势对齐效果在严格评估下不成立,即使专用模型也未能保持优势;3)领域必须采用解耦的动作质量与多模态对齐评估,以实现准确基准测试。为推动标准化并促进新评估研究,我们公开发布五小时合成动作、超过750个渲染视频刺激材料,以及16,000份配对人类偏好投票,并提供开源渲染脚本。
原文摘要 · Abstract (English)
We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed experimental setups. This leads to a situation where it is impossible to know how different methods compare, or what the state of the art is. In order to address common shortcomings of evaluation design, and to standardise future user studies in gesture-generation works, we introduce a detailed human evaluation protocol for the widely-used BEAT2 motion-capture dataset. Using this protocol, we conduct large-scale crowdsourced evaluation to rank six recent gesture-generation models -- each trained by its original authors -- across two key evaluation dimensions: motion realism and speech-gesture alignment. Our results show that 1) motion realism has become a saturated evaluation measure on the BEAT2 dataset, with older models performing on par with more recent approaches; 2) previous findings of high speech-gesture alignment do not hold up under rigorous evaluation, even for specialised models; and 3) the field must adopt disentangled assessments of motion quality and multimodal alignment for accurate benchmarking in order to make progress. To drive standardisation and enable new evaluation research, we release five hours of synthetic motion from the benchmarked models; over 750 rendered video stimuli from the user studies -- enabling new evaluations without requiring model reimplementation -- alongside our open-source rendering script, and 16,000 pairwise human preference votes collected for our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。