arXiv:2607.06843cs.CV2026-07中稿 · ECCV

通过优化初始噪声提升文本到动作生成的语义一致性。

Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation

论文配图:Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
图 1 · 摘自论文原文
  • 从随机噪声中筛选并优化出与文本对齐的'胜出噪声票'。
  • 在HumanML3D上提升文本-动作匹配度,且无需重新训练模型。
  • 适用于动作风格化、空间约束等场景,兼容多种基础模型。

基于扩散模型的文本到动作生成虽能合成逼真人体动作,但常出现语义偏离输入文本的问题。由于动作具有时序性,尤其在组合式和长时序列中,需保证多个动作片段间语义一致且轨迹平滑。我们提出,初始噪声在保持一致性中起关键作用:在高斯噪声空间中,某些特定噪声(即'胜出噪声票')蕴含潜在结构,可引导去噪过程偏向特定动作语义,即使在空提示下也如此。本文提出无训练、模型无关的WINRO框架,通过选择并优化此类噪声来改善文本-动作对齐。WINRO将随机噪声映射至空提示下的动作特征,为给定文本检索最匹配的噪声,并通过KL正则化目标减少残余语义差距,同时保留高斯先验。可选的LoRA适配器将优化过程简化为一次前向传播。WINRO在不重训的前提下,显著提升MDM与MotionLCM在HumanML3D上的文本-动作保真度,在MTT基准上增强时序鲁棒性,并可推广至动作风格化与空间约束满足等应用。

原文摘要 · Abstract (English)

Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in compositional and long-duration sequences that require semantic consistency across multiple action segments and smooth kinematic transitions throughout the trajectory. We posit that the initial noise is central to this consistency: within the Gaussian noise space, certain instances, i.e. winning noise tickets, carry latent structure that biases denoising toward particular motion semantics, even under null prompts. We propose WInning Noise Retrieval and Optimization (WINRO), a training-free, model-agnostic framework that improves text-motion alignment by selecting and refining such tickets before diffusion sampling. WINRO maps random noises to motion features generated under null prompts, retrieves the best-aligned noise for a given text, and refines it via a KL-regularized objective that reduces the residual semantic gap while preserving the Gaussian prior. An optional LoRA-based adapter amortizes this refinement into a single forward pass. WINRO consistently improves text-motion fidelity across different base models, MDM and MotionLCM, on HumanML3D without retraining, improves temporal robustness on the MTT benchmark, and generalizes to applications such as motion stylization and spatial constraint satisfaction.

动作生成扩散模型噪声优化文本对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。