用细粒度语义匹配生成罕见文本对应的自然动作
MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction
- 通过时间片段的巴赞法交互,精准衡量文本与动作片段的语义一致性
- 在罕见提示下实现当前最优的动作检索与生成效果
- 适合需要精准控制动作细节的动画生成场景
我们提出MOST,一种基于时间片段巴赞法交互的运动扩散模型,旨在解决从稀有语言提示生成人体动作的长期挑战。先前方法因运动冗余导致粗粒度匹配和关键语义线索遗漏。我们的核心洞察是利用细粒度片段关系缓解这些问题。MOST的检索阶段首次提出时间片段巴赞法交互,精确量化文本-动作在片段层面的语义一致性和匹配度,实现直接、细粒度的文本-动作片段对齐,消除普遍存在的冗余。生成阶段通过运动提示模块有效利用检索到的动作片段,生成语义一致的动作序列。大量评估表明,MOST在文本到动作的检索与生成任务中达到当前最优性能,定量与定性结果均验证其有效性,尤其在稀有提示下表现突出。
原文摘要 · Abstract (English)
We introduce MOST, a novel motion diffusion model via temporal clip Banzhaf interaction, aimed at addressing the persistent challenge of generating human motion from rare language prompts. While previous approaches struggle with coarse-grained matching and overlook important semantic cues due to motion redundancy, our key insight lies in leveraging fine-grained clip relationships to mitigate these issues. MOST's retrieval stage presents the first formulation of its kind - temporal clip Banzhaf interaction - which precisely quantifies textual-motion coherence at the clip level. This facilitates direct, fine-grained text-to-motion clip matching and eliminates prevalent redundancy. In the generation stage, a motion prompt module effectively utilizes retrieved motion clips to produce semantically consistent movements. Extensive evaluations confirm that MOST achieves state-of-the-art text-to-motion retrieval and generation performance by comprehensively addressing previous challenges, as demonstrated through quantitative and qualitative results highlighting its effectiveness, especially for rare prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。