arXiv:2605.18956cs.CV2026-05

让动作模型像人一样精细控制身体部位,实现精准编辑与推理。

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

论文配图:MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation
图 1 · 摘自论文原文
  • 在统一框架中同时建模动作的局部与时间特征,实现细粒度控制。
  • 提出跨粒度协同预训练策略,提升动作与语言的精准对齐。
  • 构建首个带时空修正指令的数据集,支持复杂动作推理与编辑。

现有动作-语言模型虽整合理解与生成任务,但仅在粗粒度层面运行,缺乏对身体部位的精细理解与可控性,制约动画与交互应用。根源在于模型无法聚焦动作的局部模式,且训练数据缺乏细粒度监督。为此,本文提出MotionMERGE统一框架:首先,开创性地在单一大模型中显式建模动作的部件与时间层级,实现细粒度语言引导动作控制;其次,设计了包含跨粒度对齐、时间定位、局部对齐、运动连贯性及动作驱动的思维链(CoT)推理的联合监督预训练策略;第三,构建大规模数据集MotionFineEdit,包含83.7万原子动作与14.4万复杂动作三元组,首次提供细粒度时空修正指令与动作驱动的思维链标注。实验表明,MotionMERGE在精确生成、理解与编辑方面表现卓越,并具备强大零样本泛化能力,推动模型向更精细粒度与类人推理迈进。

原文摘要 · Abstract (English)

Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and nuanced control over body parts needed for animation or interaction. This stems from fundamental issues in both the model and the data, in which the model can't focus on motion's localized pattern, and the training data lacks fine-grained supervision. To tackle this, we propose MotionMERGE, a unified framework that bridges the granularity gap. First, we pioneer the study of fine-grained languageguided motion control, including detailed understanding and localized editing, by explicitly modeling motion at part and temporal levels within a single LLM, thereby endowing the model with robust priors for precise control. Second, we design ReasoningAware Granularity-Synergy pre-training, a novel strategy that employs joint supervision for cross-granularity alignment, temporal grounding, localized alignment, motion coherency, and motion-grounded chain-of-thought (CoT) reasoning. This equips the model with fine-grained motion-language alignment, crossgranularity synergy, and explicit reasoning ability. Third, we curate MotionFineEdit, a large-scale dataset (837K atomic + 144K complex triplets) with the first fine-grained spatio-temporal corrective instructions and motion-grounded CoT annotations, establishing a new benchmark for fine-grained text-driven motion editing and motion-grounded reasoning. Extensive experiments demonstrate the capability of MotionMERGE for more precise motion generation, understanding, and editing, and compelling zero-shot generalization to other complex motion tasks. This work represents a significant step toward models that interact with motion in finer granularity and human-like reasoning.

动作编辑细粒度控制思维链多粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。