arXiv:2605.30608cs.CL2026-05

用语义锚点让手势更懂说话,提升人机交互自然度

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

论文配图:Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures
图 1 · 摘自论文原文
  • 将手势分解为可描述的运动单元,用语言抽象其动作与意图
  • 在BEAT2数据集上文本到手势检索准确率提升8.2%
  • 适合需要精准传达语义的手势生成与人机交互场景

学习语音与手势间的共享表示是协同说话手势检索、合成与理解的核心挑战,尤其对仅靠动作无法表达语义的手势。直接对齐文本与连续运动嵌入常过度关注低层运动特征,忽略语义手势的象征意义。本文提出语义运动锚点,即对手势物理形态与交流意图的语言化抽象。方法将3D手势离散化为肢体-手部运动基元,将其转化为结构化描述,并与文本对齐以提供辅助对比监督。在BEAT2数据集上,相比直接文本-运动基线,文本到手势检索的R@1提升8.2%,且在双向检索任务中优于已有方法。此外,语义锚点监督能有效检索出与语义相关的手势,而非通用动作模式。下游检索增强型手势生成实验表明,用户显著偏好本方法生成的手势,证明语义对齐的检索可提升生成手势的表意能力。

原文摘要 · Abstract (English)

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative intent is not captured by motion alone. Direct contrastive alignment between transcripts and continuous motion embeddings often overemphasizes low-level kinematics and misses the symbolic content of semantic gestures. We propose semantic motion anchors, natural-language abstractions of gesture motion capturing physical form and communicative intent. Our method discretizes 3D gestures into body-hand motion primitives, verbalizes them into structured descriptions, and grounds them in the transcript to provide auxiliary contrastive supervision. On BEAT2, our method improves text-to-gesture R@1 by 8.2% over a direct text-motion baseline and outperforms prior retrieval approaches on text to gesture and gesture to text retrieval directions. Beyond aggregate retrieval metrics, semantic motion anchor supervision helps retrieve gestures that are semantically meaningful for the spoken query, rather than defaulting to generic motion patterns. A downstream retrieval-augmented gesture generation study showed that users significantly preferred gestures retrieved by our approach over a retrieval-augmented generation baseline, demonstrating that semantically grounded retrieval translates to gestures that better convey communicative intent in downstream generation.

手势生成语义对齐多模态检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。