让手势理解说话意图,生成更自然有内涵的伴随动作。
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
- 将手势生成视为意图推理任务,引入高层沟通功能
- 在BEAT-2上达到新SOTA,动作时序对齐且语义丰富
- 适合数字人、具身智能等需要表达性动作的场景
人类说话时,手势有助于传达强调或描述概念等沟通意图。但当前共语言手势生成方法仅依赖语音或文本等表层语言线索,忽略了支撑手势的深层沟通意图,导致生成结果虽节奏同步却语义浅薄。为此,我们提出Intentional-Gesture框架,将手势生成视为基于高层沟通功能的意图推理任务。首先,通过大视觉语言模型自动标注,扩展BEAT-2数据集为InG数据集,添加手势意图摘要(即总结意图的文本句子)。其次,提出Intentional Gesture Motion Tokenizer,将这些意图注入动作标记表示中,实现既时间对齐又语义丰富的意图感知手势合成,在BEAT-2基准上取得新SOTA性能。该框架为数字人与具身智能中的表达性手势生成提供了模块化基础。
原文摘要 · Abstract (English)
When humans speak, gestures help convey communicative intentions, such as adding emphasis or describing concepts. However, current co-speech gesture generation methods rely solely on superficial linguistic cues (e.g. speech audio or text transcripts), neglecting to understand and leverage the communicative intention that underpins human gestures. This results in outputs that are rhythmically synchronized with speech but are semantically shallow. To address this gap, we introduce Intentional-Gesture, a novel framework that casts gesture generation as an intention-reasoning task grounded in high-level communicative functions. First, we curate the InG dataset by augmenting BEAT-2 with gesture-intention annotations (i.e., text sentences summarizing intentions), which are automatically annotated using large vision-language models. Next, we introduce the Intentional Gesture Motion Tokenizer to leverage these intention annotations. It injects high-level communicative functions (e.g., intentions) into tokenized motion representations to enable intention-aware gesture synthesis that are both temporally aligned and semantically meaningful, achieving new state-of-the-art performance on the BEAT-2 benchmark. Our framework offers a modular foundation for expressive gesture generation in digital humans and embodied AI. Project Page: https://andypinxinliu.github.io/Intentional-Gesture
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。