arXiv:2409.10357cs.CVcs.CL2024-09被引 1

比较2D与3D手势表示对语音驱动手势生成质量的影响

2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?

  • 用2D或3D关节坐标训练生成模型,对比生成效果
  • 3D训练生成的动作用客观指标和用户评价更优
  • 适合研究人机交互中手势生成的学者参考

对话手势是交流的基础。近年来深度学习技术推动了具身对话代理中逼真、同步的对话手势生成。通过人体姿态检测技术从YouTube等平台获取的“真实场景”数据集提供了与语音对齐的2D骨骼序列,为研究提供可行方案。同时,提升模型可将这些2D序列转换为3D手势数据库。然而,从2D提取的3D姿态本质上是真实姿态的近似,真实姿态仍存在于2D域中。这一差异引发关键问题:手势表示维度如何影响生成动作的质量?据我们所知,该问题尚未被充分探索。本研究考察使用2D或3D关节坐标作为训练数据对语音到手势深度生成模型性能的影响。我们采用提升模型将生成的2D姿态序列转换为3D,并评估直接在3D中生成的手势与先在2D生成再转为3D的手势之间的差异。我们使用手势生成领域常用指标进行客观评估,并开展用户研究以定性比较不同方法。

原文摘要 · Abstract (English)

Co-speech gestures are fundamental for communication. The advent of recent deep learning techniques has facilitated the creation of lifelike, synchronous co-speech gestures for Embodied Conversational Agents. "In-the-wild" datasets, aggregating video content from platforms like YouTube via human pose detection technologies, provide a feasible solution by offering 2D skeletal sequences aligned with speech. Concurrent developments in lifting models enable the conversion of these 2D sequences into 3D gesture databases. However, it is important to note that the 3D poses estimated from the 2D extracted poses are, in essence, approximations of the ground-truth, which remains in the 2D domain. This distinction raises questions about the impact of gesture representation dimensionality on the quality of generated motions - a topic that, to our knowledge, remains largely unexplored. Our study examines the effect of using either 2D or 3D joint coordinates as training data on the performance of speech-to-gesture deep generative models. We employ a lifting model for converting generated 2D pose sequences into 3D and assess how gestures created directly in 3D stack up against those initially generated in 2D and then converted to 3D. We perform an objective evaluation using widely used metrics in the gesture generation field as well as a user study to qualitatively evaluate the different approaches.

手势生成3D姿态语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。