提出可编辑的动态场景高分辨率表示,支持2D操作与3D布局统一建模。
Neural Atlas Graphs for Dynamic Scene Decomposition and Editing
- 用视图依赖的神经拼贴图作为图节点,融合2D编辑与3D空间关系
- 在Waymo数据集上比现有方法提升5 dB PSNR,实现高质量环境编辑
- 适用于驾驶以外场景,动物人类视频编辑也优于主流基线7 dB PSNR
学习可编辑的高分辨率动态场景表征是自动驾驶到创意编辑等领域的开放问题。当前方法在可编辑性与场景复杂度之间权衡:神经拼贴图将动态场景分解为前景与背景两层变形图像,支持2D编辑但难以处理多重遮挡;而场景图模型利用自动驾驶数据中的掩码与边界框捕捉复杂3D空间关系,但其隐式体素节点难以实现视角一致编辑。本文提出神经拼贴图图(NAGs),每个图节点为视图依赖的神经拼贴图,兼顾2D外观编辑与3D元素排序定位。测试时实时拟合,已在Waymo Open Dataset上达到最新量化指标,相比现有方法提升5 dB PSNR;并实现高分辨率、高质量环境编辑,可生成带有新背景和车辆外观修改的反事实驾驶场景。该方法还泛化至非驾驶场景,在DAVIS视频数据集上相较近期抠像与视频编辑基线提升超7 dB PSNR,覆盖人与动物主导的多样化场景。
原文摘要 · Abstract (English)
Learning editable high-resolution scene representations for dynamic scenes is an open problem with applications across the domains from autonomous driving to creative editing - the most successful approaches today make a trade-off between editability and supporting scene complexity: neural atlases represent dynamic scenes as two deforming image layers, foreground and background, which are editable in 2D, but break down when multiple objects occlude and interact. In contrast, scene graph models make use of annotated data such as masks and bounding boxes from autonomous-driving datasets to capture complex 3D spatial relationships, but their implicit volumetric node representations are challenging to edit view-consistently. We propose Neural Atlas Graphs (NAGs), a hybrid high-resolution scene representation, where every graph node is a view-dependent neural atlas, facilitating both 2D appearance editing and 3D ordering and positioning of scene elements. Fit at test-time, NAGs achieve state-of-the-art quantitative results on the Waymo Open Dataset - by 5 dB PSNR increase compared to existing methods - and make environmental editing possible in high resolution and visual quality - creating counterfactual driving scenarios with new backgrounds and edited vehicle appearance. We find that the method also generalizes beyond driving scenes and compares favorably - by more than 7 dB in PSNR - to recent matting and video editing baselines on the DAVIS video dataset with a diverse set of human and animal-centric scenes. Project Page: https://princeton-computational-imaging.github.io/nag/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。