arXiv:2410.10780cs.CV2024-10ICCV被引 41

提出首个可控的掩码动作生成方法,实现高精度与高质量动作合成的统一。

MaskControl: Spatio-Temporal Control for Masked Motion Synthesis

  • 通过逻辑正则化与优化,实现训练与推理时的动作精确控制。
  • 动作质量提升(FID降低约77%),控制误差降至0.91(对比1.08)。
  • 支持任意关节任意帧、身体部位时间线等多样化控制场景。

近期动作扩散模型已实现空间可控的文本到动作生成,但难以同时保证高精度控制与高质量生成。为此,我们提出MaskControl,首个为生成式掩码动作模型引入可控性的方法。该方法包含两项关键创新:一是训练时使用逻辑正则化,隐式扰动逻辑值以对齐动作标记分布与受控关节点位置,同时正则化分类标记预测以保障生成保真度;二是推理时采用逻辑优化,显式调整预测逻辑值,直接重塑标记分布,强制生成动作精准对齐受控关节点。此外,引入可微期望采样(DES)解决逻辑正则化与优化中非可微分布采样的问题。大量实验表明,MaskControl优于现有最优方法,在动作质量(FID下降约77%)和控制精度(平均误差0.91,对比1.08)上均有显著提升。同时,该方法支持任意关节任意帧控制、身体部位时间线控制及零样本目标控制等多样应用。视频演示见https://www.ekkasit.com/ControlMM-page/

原文摘要 · Abstract (English)

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these challenges, we propose MaskControl, the first approach to introduce controllability to the generative masked motion model. Our approach introduces two key innovations. First, \textit{Logits Regularizer} implicitly perturbs logits at training time to align the distribution of motion tokens with the controlled joint positions, while regularizing the categorical token prediction to ensure high-fidelity generation. Second, \textit{Logit Optimization} explicitly optimizes the predicted logits during inference time, directly reshaping the token distribution that forces the generated motion to accurately align with the controlled joint positions. Moreover, we introduce \textit{Differentiable Expectation Sampling (DES)} to combat the non-differential distribution sampling process encountered by logits regularizer and optimization. Extensive experiments demonstrate that MaskControl outperforms state-of-the-art methods, achieving superior motion quality (FID decreases by ~77\%) and higher control precision (average error 0.91 vs. 1.08). Additionally, MaskControl enables diverse applications, including any-joint-any-frame control, body-part timeline control, and zero-shot objective control. Video visualization can be found at https://www.ekkasit.com/ControlMM-page/

动作生成扩散模型控制生成姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。