arXiv:2508.13043cs.CV2025-08中稿 · the IEEE Internati…被引 1

智能引导用户拍摄更均匀的多视角图像,提升3D重建质量。

IntelliCap: Intelligent Guidance for Consistent View Sampling

  • 用视觉语言模型识别关键物体并生成球形指引
  • 实测在真实场景中优于传统采样策略
  • 适合需要高质量3D重建的摄影与扫描人员

从图像进行新颖的视图合成(如3D高斯溅射)已取得显著进展,渲染保真度和速度已能满足虚拟现实等高要求应用。然而,如何辅助人类采集输入图像这一问题却关注较少。高质量视图合成需均匀且密集的视角采样,但人类操作者常因匆忙、急躁或缺乏对场景结构和摄影过程的理解而难以满足要求。现有指导方法多针对单一物体,或忽略视图依赖的材质特性。本文提出一种多尺度的场景化可视化引导技术。在扫描过程中,该方法通过语义分割与类别识别,利用视觉-语言模型对重要物体排序,并在其周围生成球形代理以引导用户。实验表明,在真实场景中,该方法性能优于传统采样策略。

原文摘要 · Abstract (English)

Novel view synthesis from images, for example, with 3D Gaussian splatting, has made great progress. Rendering fidelity and speed are now ready even for demanding virtual reality applications. However, the problem of assisting humans in collecting the input images for these rendering algorithms has received much less attention. High-quality view synthesis requires uniform and dense view sampling. Unfortunately, these requirements are not easily addressed by human camera operators, who are in a hurry, impatient, or lack understanding of the scene structure and the photographic process. Existing approaches to guide humans during image acquisition concentrate on single objects or neglect view-dependent material characteristics. We propose a novel situated visualization technique for scanning at multiple scales. During the scanning of a scene, our method identifies important objects that need extended image coverage to properly represent view-dependent appearance. To this end, we leverage semantic segmentation and category identification, ranked by a vision-language model. Spherical proxies are generated around highly ranked objects to guide the user during scanning. Our results show superior performance in real scenes compared to conventional view sampling strategies.

3D重建智能引导图像采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。