arXiv:2412.06786cs.CV2024-12CVPR被引 36

用检索增强生成技术,让动作合成更懂语义。

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

  • 从语义知识库中检索动作示例,注入扩散模型生成过程。
  • 无需训练,通过推理时的引导机制实现语义控制。
  • 适合需要自然语义动作的虚拟角色与交互系统。

非语言交流常包含语义丰富的手势,有助于传达话语含义。现有神经网络系统虽能生成节奏性击打手势,却难以生成具有语义意义的动作。为此,我们提出 RAG-Gesture,一种基于扩散模型的姿势生成方法,利用检索增强生成(RAG)技术生成自然且富含语义的手势。该神经显式生成方法依托可解释的语言知识,通过显式领域知识从共言手势数据库中检索示例动作。检索后,借助 DDIM 反演与推理时的检索引导,将这些语义示例注入扩散生成流程,无需任何训练。此外,我们设计了一种引导控制范式,允许用户调节每次检索插入对生成序列的影响程度。对比评估表明,本方法在近期手势生成方法中表现优异。建议读者访问项目页面查看结果。

原文摘要 · Abstract (English)

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically meaningful gestures. Therefore, we present RAG-Gesture, a diffusion-based gesture generation approach that leverages Retrieval Augmented Generation (RAG) to produce natural-looking and semantically rich gestures. Our neuro-explicit gesture generation approach is designed to produce semantic gestures grounded in interpretable linguistic knowledge. We achieve this by using explicit domain knowledge to retrieve exemplar motions from a database of co-speech gestures. Once retrieved, we then inject these semantic exemplar gestures into our diffusion-based gesture generation pipeline using DDIM inversion and retrieval guidance at the inference time without any need of training. Further, we propose a control paradigm for guidance, that allows the users to modulate the amount of influence each retrieval insertion has over the generated sequence. Our comparative evaluations demonstrate the validity of our approach against recent gesture generation approaches. The reader is urged to explore the results on our project page.

手势生成扩散模型RAG语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。