arXiv:2510.17617cs.HCcs.CV2025-10

让虚拟人根据语言和图像生成有语义的肢体动作,提升沟通效果。

ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input

  • 结合语言与图像输入,零样本生成语义手势
  • 用户测试显示手势显著提升对物体属性的理解准确率
  • 适合需要自然互动的虚拟助手、数字人等场景

人类交流融合了言语与表情性非语言线索,如手部动作,承担多种交际功能。当前生成式手势方法仅限于伴随语调重复出现的节拍动作,无法传递语义信息。本文解决共语手势合成的核心挑战:生成与言语语义一致的象形或指示性手势。此类手势不能仅从语言中推断,因语言本身缺乏手势所承载的视觉意义。为此,我们提出一种零样本系统,基于语言输入并引入图像信息,无需人工标注或干预。方法整合图像分析流程,提取形状、对称性、对齐等关键物体属性,并通过语义匹配模块将这些视觉细节与语音文本关联。再利用逆运动学引擎合成象形与指示性手势,同时结合自动生成的自然节拍动作,实现连贯的多模态表达。用户研究证实其有效性:在言语模糊的场景中,本系统生成的手势显著提升了参与者对物体属性的识别能力,验证了其可解释性与交际价值。尽管复杂形状仍存挑战,结果凸显上下文感知语义手势对构建生动协作型虚拟代理的重要性,为高效可靠的具身人机交互迈出关键一步。更多详情与示例视频见:https://review-anon-io.github.io/ImaGGen.github.io/

原文摘要 · Abstract (English)

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures that accompany the rhythm of speaking but do not contribute to communicating semantic meaning. This paper tackles a core challenge in co-speech gesture synthesis: generating iconic or deictic gestures that are semantically coherent with a verbal utterance. Such gestures cannot be derived from language input alone, which inherently lacks the visual meaning that is often carried autonomously by gestures. We therefore introduce a zero-shot system that generates gestures from a given language input and additionally is informed by imagistic input, without manual annotation or human intervention. Our method integrates an image analysis pipeline that extracts key object properties such as shape, symmetry, and alignment, together with a semantic matching module that links these visual details to spoken text. An inverse kinematics engine then synthesizes iconic and deictic gestures and combines them with co-generated natural beat gestures for coherent multimodal communication. A comprehensive user study demonstrates the effectiveness of our approach. In scenarios where speech alone was ambiguous, gestures generated by our system significantly improved participants' ability to identify object properties, confirming their interpretability and communicative value. While challenges remain in representing complex shapes, our results highlight the importance of context-aware semantic gestures for creating expressive and collaborative virtual agents or avatars, marking a substantial step forward towards efficient and robust, embodied human-agent interaction. More information and example videos are available here: https://review-anon-io.github.io/ImaGGen.github.io/

手势生成多模态虚拟人零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。