arXiv:2603.20169cs.CVcs.MM2026-03被引 11

用一张第一视角图和一句指令,生成连贯的智能眼镜视频。

EgoForge: Goal-Directed Egocentric World Simulator

  • 仅需单张视图+指令+可选外部视角,生成第一人称视频。
  • 在12个任务上实现语义对齐提升18.7%,运动保真度提高23.4%。
  • 适合智能眼镜、人机交互等需自然动作模拟的场景。

生成式世界模型在动态环境模拟中展现出潜力,但第一视角视频因视角快速变化、频繁手物交互及依赖潜在人类意图的目标导向行为而难以建模。现有方法或局限于手部为中心的指令合成且场景演化有限,或仅做静态视角转换而不建模动作动态,或依赖密集监督(如相机轨迹、长视频前缀、同步多相机采集等)。本文提出EgoForge,一个目标导向的第一视角世界模拟器,仅需单张第一视角图像、高层指令和可选的辅助外部视角,即可生成连贯的第一人称视频推演。为提升意图对齐与时间一致性,我们提出VideoDiffusionNFT,一种轨迹级奖励引导的精炼机制,在扩散采样过程中优化目标完成度、时间因果性、场景一致性和感知保真度。大量实验表明,EgoForge在语义对齐、几何稳定性与运动保真度方面显著优于强基线,并在真实智能眼镜实验中表现稳健。

原文摘要 · Abstract (English)

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, synchronized multicamera capture, etc. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal consistency, we propose VideoDiffusionNFT, a trajectory-level reward-guided refinement that optimizes goal completion, temporal causality, scene consistency, and perceptual fidelity during diffusion sampling. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and robust performance in real-world smart-glasses experiments.

世界模型第一视角扩散模型智能眼镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。