arXiv:2603.25178cs.CVcs.CL2026-03

首个双语文本到动作生成基准,解决跨语言动作合成难题

Bilingual Text-to-Motion Generation: A New Benchmark and Baselines

  • 用大模型辅助标注+人工校对构建双语动作数据集
  • 提出跨语言对齐机制,零样本混语输入也能生成高质量动作
  • 在新基准上性能显著超越单语模型,适合多语言动作生成研究

文本到动作生成在跨语言应用中潜力巨大,但受限于缺乏双语数据集和现有语言模型的跨语言语义理解能力不足。为此,我们提出了首个双语文本到动作基准 BiHumanML3D,通过大语言模型辅助标注与严格人工校对构建。同时,我们提出一种简单有效的基线模型 Bilingual Motion Diffusion(BiMD),其核心为跨语言对齐(CLA)机制,显式对齐不同语言的语义表征,构建稳健的条件空间,实现从双语输入(包括零样本混语)生成高质量动作。大量实验表明,BiMD 搭载 CLA 在 BiHumanML3D 上实现 FID 0.045(对比 0.169)和 R@3 82.8%(对比 80.8%),显著优于单语扩散模型与翻译基线,验证了该数据集的必要性与对齐策略的有效性。数据集与代码已公开。

原文摘要 · Abstract (English)

Text-to-motion generation holds significant potential for cross-linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross-lingual semantic understanding of existing language models. To address these gaps, we introduce BiHumanML3D, the first bilingual text-to-motion benchmark, constructed via LLM-assisted annotation and rigorous manual correction. Furthermore, we propose a simple yet effective baseline, Bilingual Motion Diffusion (BiMD), featuring Cross-Lingual Alignment (CLA). CLA explicitly aligns semantic representations across languages, creating a robust conditional space that enables high-quality motion generation from bilingual inputs, including zero-shot code-switching scenarios. Extensive experiments demonstrate that BiMD with CLA achieves an FID of 0.045 vs. 0.169 and R@3 of 82.8\% vs. 80.8\%, significantly outperforms monolingual diffusion models and translation baselines on BiHumanML3D, underscoring the critical necessity and reliability of our dataset and the effectiveness of our alignment strategy for cross-lingual motion synthesis. The dataset and code are released at \href{https://wengwanjiang.github.io/BilingualT2M-page}{https://wengwanjiang.github.io/BilingualT2M-page}

文本生成动作双语扩散模型跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。