arXiv:2505.04974cs.CV2025-05

用奖励引导对齐,让双语文本生成更精准的人体动作。

ReAlign: Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

  • 设计步级感知奖励模型,动态评估生成过程中的语义对齐度。
  • 在多语言动作数据集上,生成动作与文本对齐度提升17.3%,质量显著改善。
  • 适合做跨语言游戏/影视动作生成的研究者和开发者使用。

双语文本到动作生成能从双语输入合成3D人体动作,在游戏、影视和机器人领域具有广阔应用前景。但该任务面临两大挑战:缺乏双语动作-语言数据集,以及扩散模型中文本与动作分布不匹配导致语义不一致或动作质量差。为此,我们构建了首个双语人体动作数据集 BiHumanML3D,为该任务提供关键基准。进一步提出双语动作扩散模型 BiMD,利用跨语言对齐表示捕捉语义,实现统一双语建模。在此基础上,提出基于奖励引导的对齐方法 ReAlign,包含步级感知奖励模型与奖励引导采样策略,通过在每一步融合文本对齐模块与动作对齐模块,动态优化噪声动作,平衡概率密度与对齐性。实验表明,相比现有最优方法,本方法显著提升文本-动作对齐度(+17.3%)与动作质量。

原文摘要 · Abstract (English)

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the absence of bilingual motion-language datasets and the misalignment between text and motion distributions in diffusion models, leading to semantically inconsistent or low-quality motions. To address these challenges, we propose BiHumanML3D, a novel bilingual human motion dataset, which establishes a crucial benchmark for bilingual text-to-motion generation models. Furthermore, we propose a Bilingual Motion Diffusion model (BiMD), which leverages cross-lingual aligned representations to capture semantics, thereby achieving a unified bilingual model. Building upon this, we propose Reward-guided sampling Alignment (ReAlign) method, comprising a step-aware reward model to assess alignment quality during sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a text-aligned module for semantic consistency and a motion-aligned module for realism, refining noisy motions at each timestep to balance probability density and alignment. Experiments demonstrate that our approach significantly improves text-motion alignment and motion quality compared to existing state-of-the-art methods. Project page: https://wengwanjiang.github.io/ReAlign-page/.

动作生成双语理解扩散模型奖励机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。