arXiv:2603.22282cs.CVcs.AI2026-03中稿 · ECCV被引 4

首个统一理解生成动作、文本与图像的连续框架

UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation

  • 将动作视为与图像同等的连续模态,构建双路径联合编码结构
  • 在7项任务中达领先效果,跨模态组合任务优势显著
  • 适合多模态生成与编辑研究者,尤其关注动作建模的场景

我们提出UniMotion,据知是首个在单一架构中同时实现人体动作、自然语言与RGB图像的理解与生成的统一框架。现有统一模型仅处理有限模态子集(如动作-文本或静态姿态-图像),且主要依赖离散标记化,导致量化误差并破坏时间连续性。UniMotion通过核心原则:将动作作为与RGB同等重要的连续模态,克服上述局限。其设计了新型跨模态对齐动作变分自编码器(CMA-VAE)和对称双路嵌入器,在共享大语言模型主干中构建动作与RGB的并行连续路径。为在推理时不需图像的情况下注入视觉语义先验,提出双后验KL对齐(DPA),将融合视觉信息的编码器更丰富的后验知识蒸馏至仅含动作的编码器。针对文本监督稀疏导致的动作路径冷启动问题,进一步提出潜变量重建对齐(LRA),一种自监督预训练策略,利用密集动作潜变量作为明确条件,协同校准嵌入器、主干网络与流动头,建立稳定的动作感知基础以支撑下游所有任务。UniMotion在涵盖任意模态间理解、生成与编辑的七项任务中均达到当前最优性能,尤其在跨模态组合任务上表现突出。

原文摘要 · Abstract (English)

We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture. Existing unified models handle only restricted modality subsets (e.g., Motion-Text or static Pose-Image) and predominantly rely on discrete tokenization, which introduces quantization errors and disrupts temporal continuity. UniMotion overcomes both limitations through a core principle: treating motion as a first-class continuous modality on equal footing with RGB. A novel Cross-Modal Aligned Motion VAE (CMA-VAE) and symmetric dual-path embedders construct parallel continuous pathways for Motion and RGB within a shared LLM backbone. To inject visual-semantic priors into motion representations without requiring images at inference, we propose Dual-Posterior KL Alignment (DPA), which distills a vision-fused encoder's richer posterior into the motion-only encoder. To address the cold-start problem -- where text supervision alone is too sparse to calibrate the newly introduced motion pathway -- we further propose Latent Reconstruction Alignment (LRA), a self-supervised pre-training strategy that uses dense motion latents as unambiguous conditions to co-calibrate the embedder, backbone, and flow head, establishing a stable motion-aware foundation for all downstream tasks. UniMotion achieves state-of-the-art performance across seven tasks spanning any-to-any understanding, generation, and editing among the three modalities, with especially strong advantages on cross-modal compositional tasks.

动作生成多模态统一框架连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。