arXiv:2512.10617cs.CV2025-12被引 1

用语言生成任意物体运动轨迹,实现精准跨模态对齐。

Lang2Motion: Bridging Language and Motion through Joint Embedding Spaces

  • 通过联合嵌入空间对齐语言与运动流形,利用双监督训练。
  • 文本到轨迹检索达34.2%召回率,运动精度提升33%-52%。
  • 支持风格迁移与语义插值,适用于多类物体运动生成。

我们提出Lang2Motion,一种通过联合嵌入空间对齐运动流形的语言引导点轨迹生成框架。不同于以往聚焦人体动作或视频合成的工作,本方法基于真实视频中的点追踪提取运动,生成任意物体的显式轨迹。基于Transformer的自编码器通过双重监督学习轨迹表示:文本运动描述与渲染轨迹可视化,二者均通过冻结的CLIP编码器映射。在文本到轨迹检索任务中,模型达到34.2% Recall@1,优于视频基线12.5个百分点;相比视频生成基线,运动精度提升33%-52%(ADE从18.3-25.3降至12.4)。尽管仅在多样物体运动上训练,仍实现88.3%的顶1准确率用于人体动作识别,展现优异跨域迁移能力。该框架支持风格迁移、语义插值及潜在空间编辑,得益于与CLIP对齐的轨迹表示。

原文摘要 · Abstract (English)

We present Lang2Motion, a framework for language-guided point trajectory generation by aligning motion manifolds with joint embedding spaces. Unlike prior work focusing on human motion or video synthesis, we generate explicit trajectories for arbitrary objects using motion extracted from real-world videos via point tracking. Our transformer-based auto-encoder learns trajectory representations through dual supervision: textual motion descriptions and rendered trajectory visualizations, both mapped through CLIP's frozen encoders. Lang2Motion achieves 34.2% Recall@1 on text-to-trajectory retrieval, outperforming video-based methods by 12.5 points, and improves motion accuracy by 33-52% (12.4 ADE vs 18.3-25.3) compared to video generation baselines. We demonstrate 88.3% Top-1 accuracy on human action recognition despite training only on diverse object motions, showing effective transfer across motion domains. Lang2Motion supports style transfer, semantic interpolation, and latent-space editing through CLIP-aligned trajectory representations.

轨迹生成跨模态语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。