arXiv:2604.23264cs.CV2026-04中稿 · CVPR被引 3

分层流匹配生成精准对齐文本的3D人体动作。

MotionHiFlow: Text-to-motion via hierarchical flow matching

论文配图:MotionHiFlow: Text-to-motion via hierarchical flow matching
图 1 · 摘自论文原文
  • 分层次建模运动:从粗到细逐步生成动作流。
  • 在HumanML3D和KIT-ML上达到当前最佳性能。
  • 适合需要精细动作对齐的研究者使用。

文本到动作生成旨在生成与输入文本紧密对齐且物理合理、细节丰富的3D人体动作。尽管现有方法能生成复杂自然的动作,但通常仅在一个时间尺度上操作,限制了语义对齐与时间连贯性。受人类认知系统中复杂动作以分层概念化而非单一时间尺度的启发,我们提出MotionHiFlow,一种分层流匹配框架,通过构建从低到高时间尺度的流路径逐步生成动作。低尺度流捕捉高层语义和粗略动作结构,高尺度流细化时间细节。为连接多尺度流,引入新颖的跨尺度过渡过程,确保连续性并保持噪声一致性。此外,通过集成文本-动作扩散Transformer和拓扑感知动作变分自编码器,利用关节感知位置编码和骨骼拓扑,显式建模关节间结构依赖,实现精确语义对齐与细微动作细节。在HumanML3D和KIT-ML基准上的大量实验表明,该方法达到当前最优性能,消融实验验证了分层设计及关键组件的有效性。代码已开源。

原文摘要 · Abstract (English)

Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which limits both semantic alignment and temporal coherence. Inspired by the fact that complex motions are conceptualized hierarchically rather than at a single temporal scale in the human cognitive system, we propose \textit{MotionHiFlow}, a hierarchical flow matching framework to generate motion progressively by constructing flow path from low to high temporal scales. The flows at lower scales capture high-level semantics and coarse motion structures, while flows at higher scales refine temporal details. To link the flows across scales, we introduce a novel cross-scale transition process, ensuring continuity and preserving noise consistency. Furthermore, by integrating a Text-Motion Diffusion Transformer and a topology-aware Motion VAE, MotionHiFlow explicitly models structural dependencies among joints via joint-aware positional encoding and skeletal topology, enabling precise semantic alignment alongside fine-grained motion details. Extensive experiments on HumanML3D and KIT-ML benchmarks demonstrate state-of-the-art performance, with ablation studies confirming the effectiveness of the hierarchical design and key components. Code is available at https://github.com/ai-lh/MotionHiFlow.

动作生成分层建模扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。