arXiv:2512.16924cs.CV2025-12被引 6

用文本+轨迹+参考图,让用户自由操控多智能体交互的动态场景生成。

The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

  • 融合文本语义、运动轨迹和参考图像,实现对物体身份与行为的精准控制。
  • 生成视频在时间上连贯,即使物体短暂消失也能保持身份一致性。
  • 适合需要复杂交互模拟的研究者或内容创作者使用。

我们提出 WorldCanvas 框架,通过结合文本、轨迹和参考图像,实现用户可控的丰富世界事件生成。与仅依赖文本的方法及现有轨迹控制的图像到视频方法不同,该多模态方法将轨迹(编码运动、时序与可见性)与自然语言(表达语义意图)及参考图像(提供物体视觉锚定)相结合,可生成包含多智能体交互、物体进出、参考引导外观变化以及反直觉事件的连贯视频。生成结果不仅具备时间一致性,还表现出涌现的一致性,即使物体临时消失,其身份与场景仍能保持稳定。该框架使世界模型从被动预测转向可交互、由用户塑造的仿真系统。项目主页见:https://worldcanvas.github.io/。

原文摘要 · Abstract (English)

We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-controlled image-to-video methods, our multimodal approach combines trajectories -- encoding motion, timing, and visibility -- with natural language for semantic intent and reference images for visual grounding of object identity, enabling the generation of coherent, controllable events that include multi-agent interactions, object entry/exit, reference-guided appearance and counterintuitive events. The resulting videos demonstrate not only temporal coherence but also emergent consistency, preserving object identity and scene despite temporary disappearance. By supporting expressive world events generation, WorldCanvas advances world models from passive predictors to interactive, user-shaped simulators. Our project page is available at: https://worldcanvas.github.io/.

可控生成多模态动态仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。