arXiv:2503.02230cs.CV2025-03AAAI被引 5

用密集视图的语义信息提升稀疏输入下的神经辐射场渲染质量

Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views

  • 利用密集新视角的语义图作为增强数据,分监督与特征两级提供引导
  • 在仅6张输入图像下仍保持高质量渲染,显著优于现有方法
  • 适合做稀疏输入场景下高保真三维重建的研究者参考

神经辐射场(NeRF)在逼真新视角生成方面表现卓越,但通常需要密集输入,稀疏输入时渲染质量会大幅下降。本文提出利用密集新视角生成的语义信息作为更鲁棒的增强数据,而非原始RGB图像。所提方法通过双层语义引导机制:监督级采用双向验证模块判断每个语义标签的有效性;特征级引入可学习代码本,通过注意力机制为每点查询语义相关特征以生成预测。该语义引导嵌入自优化流程中,并构建了一个更具挑战性的稀疏输入室内基准,输入数量最少仅6张。实验表明该方法有效,在多种指标上均优于现有方法。

原文摘要 · Abstract (English)

Neural Radiance Fields (NeRF) have shown remarkable capabilities for photorealistic novel view synthesis. One major deficiency of NeRF is that dense inputs are typically required, and the rendering quality will drop drastically given sparse inputs. In this paper, we highlight the effectiveness of rendered semantics from dense novel views, and show that rendered semantics can be treated as a more robust form of augmented data than rendered RGB. Our method enhances NeRF's performance by incorporating guidance derived from the rendered semantics. The rendered semantic guidance encompasses two levels: the supervision level and the feature level. The supervision-level guidance incorporates a bi-directional verification module that decides the validity of each rendered semantic label, while the feature-level guidance integrates a learnable codebook that encodes semantic-aware information, which is queried by each point via the attention mechanism to obtain semantic-relevant predictions. The overall semantic guidance is embedded into a self-improved pipeline. We also introduce a more challenging sparse-input indoor benchmark, where the number of inputs is limited to as few as 6. Experiments demonstrate the effectiveness of our method and it exhibits superior performance compared to existing approaches.

神经辐射场稀疏输入语义引导三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。