arXiv:2511.22098cs.CV2025-11中稿 · ECCV被引 8

实现第一人称与第三人称视频视角双向生成,提升沉浸式交互体验。

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

  • 基于上下文学习框架,融合视角对齐与协同定位编码。
  • 在EgoExo-8K数据集上达成最佳视角同步与角色一致性表现。
  • 适合游戏开发与具身智能场景的多视角视频生成需求。

近期视频世界模型的发展使得自由导航的交互环境成为可能,第一人称(egocentric)与第三人称(exocentric)视角间的转换日益重要。然而,现有研究仅关注单向的第三人称到第一人称转换,忽略了参考引导的第三人称视角合成能力,而该能力对游戏和具身智能应用至关重要。为此,我们提出WorldWander,一种专为视频生成中第一人称与第三人称世界间转换设计的上下文学习框架。基于先进的视频扩散变换器,WorldWander集成(i)上下文视角对齐与(ii)协同位置编码,以建模跨视角同步与角色一致性。为支持该任务,我们构建了EgoExo-8K数据集,包含来自合成与真实场景的同步第一人称-第三人称三元组,具有动态性与场景丰富性。实验表明,WorldWander在视角同步、角色一致性和泛化能力方面均达到最优,为第一人称-第三人称视频转换设立了新基准。

原文摘要 · Abstract (English)

Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However, existing studies focus on unidirectional exocentric-to-egocentric translation, overlooking reference-guided exocentric perspective synthesis. This capability is crucial for gaming and embodied AI applications. Motivated by this, we present WorldWander, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. Building upon advanced video diffusion transformers, WorldWander integrates (i) In-Context Perspective Alignment and (ii) Collaborative Position Encoding to model cross-view synchronization and character consistency. To support our task, we curate EgoExo-8K, a dynamic and scene-rich dataset containing synchronized egocentric-exocentric triplets from both synthetic and real-world scenarios. Experiments demonstrate that WorldWander achieves superior perspective synchronization, character consistency, and generalization, setting a new benchmark for egocentric-exocentric video translation.

视频生成视角转换具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。