用音频指令生成符合语义的人体动作,更自然直观。
Semantics-Aware Human Motion Generation from Audio Instructions
- 用掩码生成式Transformer+记忆检索注意力处理长音频输入
- 在新构建数据集上实现音频与动作语义高度对齐
- 适合需要自然交互的虚拟人、机器人动作生成场景
近年来,交互技术的发展凸显了音频信号在语义编码中的重要性。本文提出一项新任务:以音频信号为条件生成与音频语义一致的动作。相比文本交互,音频提供更自然直观的沟通方式。然而,现有方法多关注音乐或语音节奏匹配,导致音频语义与生成动作关联较弱。本文提出一种端到端框架,采用掩码生成式Transformer,并引入记忆检索注意力模块,有效处理稀疏且长时的音频输入。同时,通过将描述转换为对话风格并生成不同说话人身份的音频,丰富了现有数据集。实验表明,该框架在有效性与效率上均表现优异,证明音频指令可传递与文本相当的语义信息,且具备更强的实用性与用户友好性。
原文摘要 · Abstract (English)
Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the semantics of the audio. Unlike text-based interactions, audio provides a more natural and intuitive communication method. However, existing methods typically focus on matching motions with music or speech rhythms, which often results in a weak connection between the semantics of the audio and generated motions. We propose an end-to-end framework using a masked generative transformer, enhanced by a memory-retrieval attention module to handle sparse and lengthy audio inputs. Additionally, we enrich existing datasets by converting descriptions into conversational style and generating corresponding audio with varied speaker identities. Experiments demonstrate the effectiveness and efficiency of the proposed framework, demonstrating that audio instructions can convey semantics similar to text while providing more practical and user-friendly interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。