用大模型解析说话意图,生成语义连贯的手势。
SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain
- 通过意图链分解说话内容,生成结构化手势标签
- 手势标注准确率达50.2%,单次推理仅需0.4秒
- 可解释的推理路径,适合需要自然互动的应用
同步语音手势生成能提升人机交互的真实感,但生成具有语义意义的手势仍是难题。本文提出SARGes框架,利用大语言模型(LLMs)解析语音内容并生成可靠语义手势标签,进而指导有意义的同步语音手势合成。首先,构建了全面的同步语音手势图谱,并设计基于LLM的意图链推理机制,将手势语义按图谱标准系统分解为结构化推理步骤,有效引导模型生成上下文感知的手势标签。随后,构建了带有意图链标注的文本到手势标签数据集,并训练轻量级手势标签生成模型,从而指导生成可信且语义连贯的同步语音手势。实验表明,SARGes在手势标注上达到50.2%的准确率,支持高效单次推理(0.4秒)。该方法为语义手势合成提供了可解释的意图推理路径。
原文摘要 · Abstract (English)
Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGes, a novel framework that leverages large language models (LLMs) to parse speech content and generate reliable semantic gesture labels, which subsequently guide the synthesis of meaningful co-speech gestures.First, we constructed a comprehensive co-speech gesture ethogram and developed an LLM-based intent chain reasoning mechanism that systematically parses and decomposes gesture semantics into structured inference steps following ethogram criteria, effectively guiding LLMs to generate context-aware gesture labels. Subsequently, we constructed an intent chain-annotated text-to-gesture label dataset and trained a lightweight gesture label generation model, which then guides the generation of credible and semantically coherent co-speech gestures. Experimental results demonstrate that SARGes achieves highly semantically-aligned gesture labeling (50.2% accuracy) with efficient single-pass inference (0.4 seconds). The proposed method provides an interpretable intent reasoning pathway for semantic gesture synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。