arXiv:2412.17561cs.CV2024-12AAAI被引 6

用隐式神经场提升室内场景生成的真实感

S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural Field

  • 分离场景布局与物体细节关系,通过隐式神经场融合
  • 在3D-FRONT数据集上实现当前最优的场景生成效果
  • 适合关注真实场景布局与风格一致性的生成研究者

基于学习的方法在3D室内场景合成(ISS)中日益流行,性能优于传统优化方法。但现有方法通常依赖简化的显式场景表示,忽略细节信息,且缺乏场景内多模态关系的引导,导致生成的场景在物体布局和风格上不够真实。本文提出场景隐式神经场(S-INF),通过解耦场景布局关系与物体细节关系,并利用隐式神经场(INFs)进行融合,学习多模态关系。S-INF通过专门学习场景布局关系并投影至模型,生成更真实的场景布局;同时借助可微渲染捕捉密集的物体关系,确保物体间风格一致性。在基准数据集3D-FRONT上的大量实验表明,该方法在不同类型的室内场景合成任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Learning-based methods have become increasingly popular in 3D indoor scene synthesis (ISS), showing superior performance over traditional optimization-based approaches. These learning-based methods typically model distributions on simple yet explicit scene representations using generative models. However, due to the oversimplified explicit representations that overlook detailed information and the lack of guidance from multimodal relationships within the scene, most learning-based methods struggle to generate indoor scenes with realistic object arrangements and styles. In this paper, we introduce a new method, Scene Implicit Neural Field (S-INF), for indoor scene synthesis, aiming to learn meaningful representations of multimodal relationships, to enhance the realism of indoor scene synthesis. S-INF assumes that the scene layout is often related to the object-detailed information. It disentangles the multimodal relationships into scene layout relationships and detailed object relationships, fusing them later through implicit neural fields (INFs). By learning specialized scene layout relationships and projecting them into S-INF, we achieve a realistic generation of scene layout. Additionally, S-INF captures dense and detailed object relationships through differentiable rendering, ensuring stylistic consistency across objects. Through extensive experiments on the benchmark 3D-FRONT dataset, we demonstrate that our method consistently achieves state-of-the-art performance under different types of ISS.

场景合成隐式表示多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。