arXiv:2411.02179cs.CVcs.GR2024-11被引 7

用生成模型快速准确估算手机AR环境光照,让虚拟物体更真实。

CleAR: Robust Context-Guided Generative Lighting Estimation for Mobile Augmented Reality

  • 分两步生成,结合环境上下文确保光照贴合真实场景。
  • 3.2秒出结果,比现有方法快110倍,材质渲染准确率提升51%-56%。
  • 适合做移动AR光影建模的开发者,尤其关注实时性与真实感的人。

高质量环境光照对实现沉浸式移动增强现实(AR)体验至关重要。然而,由于移动AR设备传感能力的限制,如摄像头视场角小、像素动态范围有限,实现视觉一致的光照估计极具挑战。近期生成式AI技术(尤其是扩散模型)可从文本或图像等提示中生成高质量图像,为高保真光照估计提供了新路径。但要有效利用生成模型,仍需解决内容质量与推理速度两大难题。本文设计并实现了名为CleAR的生成式光照估计系统,能够生成高质量、多样化的360° HDR环境图。具体而言,我们提出一种由AR环境上下文数据引导的两阶段生成流程,确保输出与物理环境的视觉上下文及色彩外观保持一致;为提升不同光照条件下的鲁棒性,还设计了实时优化模块,在AR设备上动态调整光照估计结果。通过定量与定性评估,我们发现CleAR在估计精度、延迟和鲁棒性方面均优于现有最佳方法。31名参与者评价其对多数虚拟物体的渲染效果更佳。例如,针对三种不同材质和反射特性的物体,其虚拟物体渲染准确率提升51%至56%。CleAR仅需3.2秒即可生成光照估计,较当前最优方法提速超过110倍。

原文摘要 · Abstract (English)

High-quality environment lighting is essential for creating immersive mobile augmented reality (AR) experiences. However, achieving visually coherent estimation for mobile AR is challenging due to several key limitations in AR device sensing capabilities, including low camera FoV and limited pixel dynamic ranges. Recent advancements in generative AI, which can generate high-quality images from different types of prompts, including texts and images, present a potential solution for high-quality lighting estimation. Still, to effectively use generative image diffusion models, we must address two key limitations of content quality and slow inference. In this work, we design and implement a generative lighting estimation system called CleAR that can produce high-quality, diverse environment maps in the format of 360° HDR images. Specifically, we design a two-step generation pipeline guided by AR environment context data to ensure the output aligns with the physical environment's visual context and color appearance. To improve the estimation robustness under different lighting conditions, we design a real-time refinement component to adjust lighting estimation results on AR devices. Through a combination of quantitative and qualitative evaluations, we show that CleAR outperforms state-of-the-art lighting estimation methods on both estimation accuracy, latency, and robustness, and is rated by 31 participants as producing better renderings for most virtual objects. For example, CleAR achieves 51% to 56% accuracy improvement on virtual object renderings across objects of three distinctive types of materials and reflective properties. CleAR produces lighting estimates of comparable or better quality in just 3.2 seconds -- over 110X faster than state-of-the-art methods.

AR光照生成模型实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。