arXiv:2603.19795cs.CV2026-03被引 1

用相位控制实现动作生成中局部肢体的精准编辑

Controllable Text-to-Motion Generation via Modular Body-Part Phase Control

  • 将肢体运动建模为可调相位信号,实现精细控制
  • 支持对动作幅度、速度和时机的可预测调节
  • 适合需要交互式修改动作的动画与虚拟角色开发

文本到动作(T2M)生成正成为动画和交互式虚拟角色的实用工具。然而,在保持整体动作连贯性的前提下修改特定身体部位仍具挑战。现有方法通常依赖复杂且高维的关节约束(如轨迹),难以实现用户友好的迭代优化。为此,我们提出模块化肢体相位控制(Modular Body-Part Phase Control),一种即插即用框架,通过紧凑的标量相位接口实现结构化、局部化编辑。我们将肢体潜空间运动通道建模为由振幅、频率、相位偏移和偏置定义的正弦相位信号,提取可解释的编码以捕捉各部位特异性动态。一个模块化的相位控制网络分支通过残差特征调制注入该信号,无缝解耦控制与生成主干。在基于扩散模型和基于流的模型上的实验表明,本方法能实现对动作幅度、速度和时机的可预测、细粒度控制,同时保持全局动作连贯性,为可控T2M生成提供了实用范式。

原文摘要 · Abstract (English)

Text-to-motion (T2M) generation is becoming a practical tool for animation and interactive avatars. However, modifying specific body parts while maintaining overall motion coherence remains challenging. Existing methods typically rely on cumbersome, high-dimensional joint constraints (e.g., trajectories), which hinder user-friendly, iterative refinement. To address this, we propose Modular Body-Part Phase Control, a plug-and-play framework enabling structured, localized editing via a compact, scalar-based phase interface. By modeling body-part latent motion channels as sinusoidal phase signals characterized by amplitude, frequency, phase shift, and offset, we extract interpretable codes that capture part-specific dynamics. A modular Phase ControlNet branch then injects this signal via residual feature modulation, seamlessly decoupling control from the generative backbone. Experiments on both diffusion- and flow-based models demonstrate that our approach provides predictable and fine-grained control over motion magnitude, speed, and timing. It preserves global motion coherence and offers a practical paradigm for controllable T2M generation. Project page: https://jixiii.github.io/bp-phase-project-page/

动作生成相位控制可编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。