用生成模型提前预测带编辑的增强现实画面,再局部修正关键区域。
SEGAR: Selective Enhancement for Generative Augmented Reality

- 基于扩散模型生成带区域编辑的未来画面,保持时间连贯性。
- 通过局部修正确保安全关键区域与真实世界一致,其他区域保留编辑效果。
- 适用于自动驾驶等场景,适合需要低延迟实时渲染的AR应用。
生成式世界模型为增强现实(AR)应用提供了有力基础:通过预测包含特定视觉修改的未来图像序列,可提前计算并缓存时序连贯的增强帧,避免实时逐帧重渲染。本文提出SEGAR框架,将基于扩散的世界模型与选择性修正阶段结合,实现这一目标。世界模型生成带有特定区域编辑的增强未来帧,同时保持其他区域不变;修正阶段则对安全关键区域进行对齐,使其符合真实世界观测,同时保留其他区域的预期增广效果。我们在驾驶场景中验证该流程,该场景语义区域结构清晰且真实反馈易获取。本工作被视为迈向生成式世界模型作为实用AR基础设施的初步探索,未来帧可预先生成、缓存,并按需选择性修正。
原文摘要 · Abstract (English)
Generative world models offer a compelling foundation for augmented-reality (AR) applications: by predicting future image sequences that incorporate deliberate visual edits, they enable temporally coherent, augmented future frames that can be computed ahead of time and cached, avoiding per-frame rendering from scratch in real time. In this work, we present SEGAR, a preliminary framework that combines a diffusion-based world model with a selective correction stage to support this vision. The world model generates augmented future frames with region-specific edits while preserving others, and the correction stage subsequently aligns safety-critical regions with real-world observations while preserving intended augmentations elsewhere. We demonstrate this pipeline in driving scenarios as a representative setting where semantic region structure is well defined and real-world feedback is readily available. We view this as an early step toward generative world models as practical AR infrastructure, where future frames can be generated, cached, and selectively corrected on demand.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。