通过引入随机不确定性,让文本生成3D人体动作更丰富多样。
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
- 用噪声信号承载多样性信息,显式建模生成过程中的不确定性。
- 在连续潜在空间中采样,实现动作生成的随机性,提升多样性。
- 在HumanML3D和KIT-ML数据集上兼顾语义一致性与动作多样性。
从文本生成3D人体动作为一项挑战性但有价值的任务,其核心在于确保文本与动作的一致性以及生成动作的多样性。尽管近期进展已实现高精度、高质量的文本到动作生成,但如何提升生成动作的多样性仍是关键难题。本文提出一种简单而有效的方法Diverse-T2M,通过在生成过程中引入不确定性,显著增强动作多样性,同时保持文本语义一致性。具体而言,我们创新性地将噪声信号作为变换器模型中多样性信息的载体,实现对不确定性的显式建模;此外,构建一个连续的潜在空间,将文本映射为非刚性的一对多表示,并引入潜在空间采样器,实现生成过程中的随机采样,从而提升输出的多样性和不确定性。在HumanML3D和KIT-ML两个基准数据集上的实验结果表明,该方法在保持文本一致性达到顶尖水平的同时,显著提升了生成动作的多样性。
原文摘要 · Abstract (English)
Generating 3D human motions from text is a challenging yet valuable task. The key aspects of this task are ensuring text-motion consistency and achieving generation diversity. Although recent advancements have enabled the generation of precise and high-quality human motions from text, achieving diversity in the generated motions remains a significant challenge. In this paper, we aim to overcome the above challenge by designing a simple yet effective text-to-motion generation method, \textit{i.e.}, Diverse-T2M. Our method introduces uncertainty into the generation process, enabling the generation of highly diverse motions while preserving the semantic consistency of the text. Specifically, we propose a novel perspective that utilizes noise signals as carriers of diversity information in transformer-based methods, facilitating a explicit modeling of uncertainty. Moreover, we construct a latent space where text is projected into a continuous representation, instead of a rigid one-to-one mapping, and integrate a latent space sampler to introduce stochastic sampling into the generation process, thereby enhancing the diversity and uncertainty of the outputs. Our results on text-to-motion generation benchmark datasets~(HumanML3D and KIT-ML) demonstrate that our method significantly enhances diversity while maintaining state-of-the-art performance in text consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。