arXiv:2509.23852cs.GRcs.MM2025-09

让对话中的手势更精准,能自动判断何时、何地、怎么比。

SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where

  • 结合语音、语言和空间信息生成有交互意图的手势
  • 在人形机器人上实现时序与空间精准的肢体互动
  • 适合需要自然交互的机器人、游戏和动画场景

对话中的伴随动作和手势常与环境互动紧密相关,如看向对话者或在恰当时机用手势指向目标。语音与语义决定手势的时机(何时)和风格(如何),而交互对象的空间位置则决定其方向性执行(何处)。现有方法要么仅依赖描述性语言生成动作,要么使用音频生成非交互手势,缺乏对交互时机和空间意图的刻画,严重限制了对话手势生成在机器人及游戏动画领域的应用。为此,我们提出端到端解决方案:首先建立一种独特数据采集方法,同步获取高精度人体运动与空间意图;随后开发基于音频、语言与空间数据的生成模型,并设计专用评估指标以衡量交互时机与空间准确性;最终在人形机器人上部署,实现丰富且上下文感知的物理交互。

原文摘要 · Abstract (English)

The accompanying actions and gestures in dialogue are often closely linked to interactions with the environment, such as looking toward the interlocutor or using gestures to point to the described target at appropriate moments. Speech and semantics guide the production of gestures by determining their timing (WHEN) and style (HOW), while the spatial locations of interactive objects dictate their directional execution (WHERE). Existing approaches either rely solely on descriptive language to generate motions or utilize audio to produce non-interactive gestures, thereby lacking the characterization of interactive timing and spatial intent. This significantly limits the applicability of conversational gesture generation, whether in robotics or in the fields of game and animation production. To address this gap, we present a full-stack solution. We first established a unique data collection method to simultaneously capture high-precision human motion and spatial intent. We then developed a generation model driven by audio, language, and spatial data, alongside dedicated metrics for evaluating interaction timing and spatial accuracy. Finally, we deployed the solution on a humanoid robot, enabling rich, context-aware physical interactions.

手势生成人机交互空间意图机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。