arXiv:2510.13208cs.CVcs.AI2025-10被引 2

让语音驱动的3D动作更自然,按身体部位注入风格

MimicParts: Part-aware Style Injection for Speech-Driven 3D Motion Generation

  • 按上肢、下肢等区域分别编码风格,捕捉局部差异
  • 语音节奏与情绪变化能精准引导各部位动作
  • 适合做角色动画、虚拟人交互等需要细腻动作的场景

从语音信号生成具有风格化的3D人体动作面临重大挑战,主要源于语音、风格与身体动作之间复杂且细微的关联。现有风格编码方法或过度简化风格多样性,或忽略局部动作风格差异(如上肢与下肢),限制了动作的真实感。此外,动作风格应随语音节奏和情感动态调整,但现有方法常忽视这一点。为此,我们提出MimicParts,一种基于部位感知风格注入与部位感知去噪网络的新框架。该框架将身体划分为不同区域,对局部运动风格进行编码,使模型能捕捉细微的区域差异。同时,部位感知注意力模块可精确引导各身体区域响应语音节奏与情感线索,确保生成动作与语音变化同步。实验结果表明,本方法在自然度与表现力方面优于现有方法,生成更具表现力的3D人体动作序列。

原文摘要 · Abstract (English)

Generating stylized 3D human motion from speech signals presents substantial challenges, primarily due to the intricate and fine-grained relationships among speech signals, individual styles, and the corresponding body movements. Current style encoding approaches either oversimplify stylistic diversity or ignore regional motion style differences (e.g., upper vs. lower body), limiting motion realism. Additionally, motion style should dynamically adapt to changes in speech rhythm and emotion, but existing methods often overlook this. To address these issues, we propose MimicParts, a novel framework designed to enhance stylized motion generation based on part-aware style injection and part-aware denoising network. It divides the body into different regions to encode localized motion styles, enabling the model to capture fine-grained regional differences. Furthermore, our part-aware attention block allows rhythm and emotion cues to guide each body region precisely, ensuring that the generated motion aligns with variations in speech rhythm and emotional state. Experimental results show that our method outperforming existing methods showcasing naturalness and expressive 3D human motion sequences.

3D动作生成语音驱动风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。