arXiv:2503.14408cs.HCcs.CL2025-03中稿 · the AAMAS 2025 con…被引 6

用大模型自动选对话手势,让虚拟人更自然互动。

Large Language Models for Virtual Human Gesture Selection

  • 利用大模型语义能力,智能匹配说话内容与手势
  • 实验证明能选出符合语境的有意义手势
  • 适合虚拟人交互、人机对话系统研发者

对话手势蕴含丰富含义,在面对面交流中显著影响听者的参与度、记忆、理解与态度。同样,它们也深刻影响人类与具身虚拟代理之间的互动。因此,选择并动画化有意义的手势成为虚拟代理设计的关键。然而,自动化手势选择面临挑战:以往方法或完全依赖数据驱动,生成的手势缺乏语境意义;或依赖人工设计,耗时且难以泛化。本文利用大语言模型的语义能力,提出一种手势选择方法,可生成恰当的对话手势。我们首先分析手势信息在GPT-4中的编码方式,再通过实验评估不同提示策略在选取语境相关手势上的效果,并实现其在虚拟代理系统中的集成,自动完成手势选择与动画,提升人机交互体验。

原文摘要 · Abstract (English)

Co-speech gestures convey a wide variety of meanings and play an important role in face-to-face human interactions. These gestures significantly influence the addressee's engagement, recall, comprehension, and attitudes toward the speaker. Similarly, they impact interactions between humans and embodied virtual agents. The process of selecting and animating meaningful gestures has thus become a key focus in the design of these agents. However, automating this gesture selection process poses a significant challenge. Prior gesture generation techniques have varied from fully automated, data-driven methods, which often struggle to produce contextually meaningful gestures, to more manual approaches that require crafting specific gesture expertise and are time-consuming and lack generalizability. In this paper, we leverage the semantic capabilities of Large Language Models to develop a gesture selection approach that suggests meaningful, appropriate co-speech gestures. We first describe how information on gestures is encoded into GPT-4. Then, we conduct a study to evaluate alternative prompting approaches for their ability to select meaningful, contextually relevant gestures and to align them appropriately with the co-speech utterance. Finally, we detail and demonstrate how this approach has been implemented within a virtual agent system, automating the selection and subsequent animation of the selected gestures for enhanced human-agent interactions.

虚拟人手势生成大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。