对比两种手势生成框架,发现语义增强未必更好
Conveying Meaning through Gestures: An Investigation into Semantic Co-Speech Gesture Generation
- 用语义增强与原始框架生成对话手势
- 原框架在训练域内概念传达更有效,增强版泛化更强
- 增强版更生动但不更像人,说明语义注入有局限
本研究探讨了两种对话手势生成框架——AQ-GT及其语义增强变体AQ-GT-a——在传达语义方面的表现及人类对生成动作的感知。基于SAGA空间交流语料库中的句子、上下文相似句和全新运动导向句,进行了以用户为中心的概念识别与自然度评估。结果揭示了语义标注与性能之间的复杂关系:未显式引入语义信息的原始AQ-GT框架在训练领域内概念传达效果更优;而AQ-GT-a在新情境中对形状与大小的表征具有更强泛化能力。尽管参与者认为AQ-GT-a生成的手势更具表现力和帮助性,但并未感觉其更像真人动作。这表明显式语义增强并不必然提升手势生成质量,其有效性高度依赖于任务场景,暗示专精与泛化之间存在权衡。
原文摘要 · Abstract (English)
This study explores two frameworks for co-speech gesture generation, AQ-GT and its semantically-augmented variant AQ-GT-a, to evaluate their ability to convey meaning through gestures and how humans perceive the resulting movements. Using sentences from the SAGA spatial communication corpus, contextually similar sentences, and novel movement-focused sentences, we conducted a user-centered evaluation of concept recognition and human-likeness. Results revealed a nuanced relationship between semantic annotations and performance. The original AQ-GT framework, lacking explicit semantic input, was surprisingly more effective at conveying concepts within its training domain. Conversely, the AQ-GT-a framework demonstrated better generalization, particularly for representing shape and size in novel contexts. While participants rated gestures from AQ-GT-a as more expressive and helpful, they did not perceive them as more human-like. These findings suggest that explicit semantic enrichment does not guarantee improved gesture generation and that its effectiveness is highly dependent on the context, indicating a potential trade-off between specialization and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。