arXiv:2602.04876cs.CV2026-02被引 7

从单图生成可交互的长期动态4D场景,物理与视觉闭环联动。

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

  • 构建物理状态与视觉表征的双向关联统一表示
  • 支持多视角监督,实现长期动作下物理合理且视觉一致的生成
  • 适合需要真实感动态场景生成的研究与应用

我们提出 PerpetualWonder,一种混合生成式模拟器,能够从单张图像生成长期、动作条件化的4D场景。现有方法在此任务上表现不佳,因其物理状态与视觉表示解耦,导致生成优化无法修正后续交互的底层物理。PerpetualWonder 通过引入首个真正的闭环系统解决该问题:采用新颖的统一表示,建立物理状态与视觉基元之间的双向链接,使生成优化可同时修正动态与外观。同时,引入稳健的更新机制,通过多视角监督消除优化歧义。实验表明,从单张图像出发,PerpetualWonder 可成功模拟复杂、多步骤的长期动作交互,在保持物理合理性的同时维持视觉一致性。

原文摘要 · Abstract (English)

We introduce PerpetualWonder, a hybrid generative simulator that enables long-horizon, action-conditioned 4D scene generation from a single image. Current works fail at this task because their physical state is decoupled from their visual representation, which prevents generative refinements to update the underlying physics for subsequent interactions. PerpetualWonder solves this by introducing the first true closed-loop system. It features a novel unified representation that creates a bidirectional link between the physical state and visual primitives, allowing generative refinements to correct both the dynamics and appearance. It also introduces a robust update mechanism that gathers supervision from multiple viewpoints to resolve optimization ambiguity. Experiments demonstrate that from a single image, PerpetualWonder can successfully simulate complex, multi-step interactions from long-horizon actions, maintaining physical plausibility and visual consistency.

4D生成动作条件闭环模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。