arXiv:2605.29219cs.CV2026-05

让机器人能听音乐、跟人跳舞,生成自然同步的桑巴舞动作。

SalsaAgent: A multimodal embodied language model for interactive dance generation

论文配图:SalsaAgent: A multimodal embodied language model for interactive dance generation
图 1 · 摘自论文原文
  • 用动作符号和关系符号扩展语言模型,实现舞步与音乐、搭档的实时互动。
  • 在多人舞蹈场景中,动作质量与音乐节奏、搭档协调性显著优于基线方法。
  • 适合做交互式虚拟角色或社交机器人的开发者参考。

人形机器人之间的互动需要双向、非语言的反应、协调与同步。为实现具有社会意识的机器人和交互式虚拟角色,我们提出 SalsaAgent,一个能够根据人类领舞者和背景音乐生成富有表现力的全身桑巴舞动作的语言模型。我们将交互建模为非语言动作符号的传递,扩展大语言模型(LLM)的词汇表,使其可处理离散动作符号、成对关系符号和音频信号。我们的贡献包括全新设计的全身动作与动作关系符号,利用自动提取的动作描述对齐动作符号进行微调,以及两阶段从符号到扩散模型的动作生成管道。主观与客观评估表明,该方法在动作质量、音乐与舞伴协调性、两人空间行为一致性方面均显著优于基线模型。

原文摘要 · Abstract (English)

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generates expressive, full-body salsa dance motions in reaction to a human leader and against a contextual music backdrop. We formulate interaction as nonverbal motion token passing, extending the vocabulary of a large language model (LLM) to process discrete motion tokens, pairwise relation tokens, and audio. Our contributions include new tokens for full-body and motion relations, LLM fine-tuning using automatically derived text descriptions of skeleton dynamics for token grounding, and a two-stage token-to-diffusion pipeline. Subjective and objective evaluations demonstrate the effectiveness of our approach in terms of motion quality, music and partner coordination, and consistent two-person spatial behavior, with significant improvements over baselines.

舞蹈生成多模态语言模型机器人互动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。