arXiv:2604.11737cs.CV2026-04被引 1

用压缩运动嵌入高效生成符合指令的长时序动作。

Learning Long-term Motion Embeddings for Efficient Kinematics Generation

论文配图:Learning Long-term Motion Embeddings for Efficient Kinematics Generation
图 1 · 摘自论文原文
  • 通过64倍时间压缩学习长期运动嵌入,实现高效建模。
  • 在文本或空间提示下生成符合目标的长序列真实动作。
  • 性能超越主流视频模型和专用任务方法,适合动作生成场景。

理解与预测运动是视觉智能的基础。尽管现代视频模型对场景动态有较强理解能力,但通过完整视频合成探索多种可能未来仍效率极低。本文通过直接操作从追踪模型获得的大规模轨迹中学习的长期运动嵌入,将场景动态建模效率提升数个数量级。该方法可高效生成满足文本提示或空间点击目标的长时序、高保真动作。为此,我们首先以64倍的时间压缩因子学习高度压缩的运动嵌入,并在此空间中训练条件流匹配模型,根据任务描述生成运动隐变量。生成的运动分布在性能上优于当前最先进的视频模型及专用任务方法。

原文摘要 · Abstract (English)

Understanding and predicting motion is a fundamental component of visual intelligence. Although modern video models exhibit strong comprehension of scene dynamics, exploring multiple possible futures through full video synthesis remains prohibitively inefficient. We model scene dynamics orders of magnitude more efficiently by directly operating on a long-term motion embedding that is learned from large-scale trajectories obtained from tracker models. This enables efficient generation of long, realistic motions that fulfill goals specified via text prompts or spatial pokes. To achieve this, we first learn a highly compressed motion embedding with a temporal compression factor of 64x. In this space, we train a conditional flow-matching model to generate motion latents conditioned on task descriptions. The resulting motion distributions outperform those of both state-of-the-art video models and specialized task-specific approaches.

动作生成运动嵌入扩散模型长时序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。