arXiv:2604.13427cs.GRcs.AI2026-04被引 1

一个模型搞定动作生成、编辑和骨骼适配,无需分步训练。

A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting

  • 用统一流模型同时处理动作生成与编辑,条件可自由切换。
  • 在SnapMoGen和Mixamo数据集上实现零样本编辑与骨骼适配。
  • 适合做动作生成、动画编辑的开发者或研究人员使用。

文本驱动的动作编辑与同构结构骨骼适配(骨骼拓扑相同但长度和静止姿态不同)传统上依赖不兼容的独立流程:编辑依赖特定生成引导,而适配则需几何后处理。本文提出统一的条件流框架,将生成、语义编辑与同构结构适配统一为一个文本与骨骼条件控制的修正流模型中的条件调控传输过程。在此框架下,编辑改变语义条件但保持骨骼结构不变,适配则改变骨骼条件但保留动作语义。该方法使FlowEdit式传输成为通用动作操控规则,而非任务专用编辑器。为实现3D关节动作的统一建模,我们设计了一个文本与骨骼条件控制的修正流变压器模型,采用逐关节标记化与显式关节自注意力捕捉空间运动学依赖。同时在关节与帧层级注入文本条件,并通过残差多条件无分类器引导平衡文本一致性与骨骼符合性。在SnapMoGen与多角色Mixamo子集上的实验表明,仅需一个训练好的模型即可支持文本到动作生成、零样本编辑及零样本同构结构适配,无需针对每项任务微调。该统一框架以单一条件动作传输模型替代原有多个分离流程,同时明确保留同拓扑适配的适用范围。

原文摘要 · Abstract (English)

Text-driven motion editing and intra-structural retargeting, where skeletons share topology but may differ in bone lengths and rest pose, are traditionally handled by fragmented pipelines with incompatible inputs and representations: editing relies on specialized generative steering, while retargeting is deferred to geometric post-processing. We present a unified conditional-flow framework that casts generation, semantic editing, and intra-structural retargeting as condition-modulated transport within one text- and skeleton-conditioned rectified-flow model. Under this formulation, editing changes the semantic condition while preserving skeletal structure, whereas retargeting changes the skeletal condition while preserving motion semantics. This makes FlowEdit-style transport a unified inference rule for motion manipulation rather than a task-specific editor. To instantiate this for articulated 3D motion, we develop a text- and skeleton-conditioned rectified-flow transformer. The model uses per-joint tokenization and explicit joint self-attention to capture spatial kinematic dependencies. We further inject text conditions at both joint and frame levels, while residual multi-condition classifier-free guidance balances text adherence and skeletal conformity. Experiments on SnapMoGen and a multi-character Mixamo subset show that one trained model supports text-to-motion generation, zero-shot editing, and zero-shot intra-structural retargeting without task-specific fine-tuning. This unified framework replaces separate pipelines with a single conditional motion transport model while keeping the same-topology retargeting scope explicit.

动作生成文本控制骨骼适配统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。