arXiv:2608.29177cs.CV2026-08中稿 · ECCV

解决动态场景下3D重建与语义分割的错位问题,实现更稳定的视觉理解。

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

论文配图:Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
图 1 · 摘自论文原文
  • 先分离动态噪声再融合特征,提升动态环境下的几何-语义一致性。
  • 仅用3/4张输入图即达22.15/23.33 dB PSNR,自监督训练下语义掩码mIoU达88.5%。
  • 揭示了图像重建与语义理解间的协同增益,适合动态3D场景研究者。

将新视角合成(NVS)与开放词汇分割(OVS)结合的3D基础模型虽强大,但依赖静态场景假设,在动态环境中导致空间特征严重错位。为此,我们提出SPAR,一种联合语义-几何编码架构,能预先显式分离瞬态动态噪声,再进行潜在空间聚合。同时引入动态区域感知的端到端训练范式,将运动估计与多视角视觉及语义学习结构耦合。该统一方法使网络能内生化解运动冲突,并从动态输入中提炼出一致且时序稳定的场景表征。在挑战性D-RE10K基准上,实验表明SPAR达到领先性能:仅用3/4张输入视图即实现22.15/23.33 dB的PSNR;自监督训练下,运动掩码预测的mIoU达88.5%。分析揭示光度重建与语义理解间存在强任务协同,语义合成学习持续提升新视角渲染的保真度。代码将在https://github.com/dmucby/SPAR公开。

原文摘要 · Abstract (English)

The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic-region-aware end-to-end training paradigm that structurally couples motion estimation with multi-view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi-view consistent, temporally stable scene representations from dynamic inputs. Extensive experiments on the challenging D-RE10K benchmark demonstrate that SPAR achieves state-of-the-art performance. Our end-to-end approach achieves exceptional novel view synthesis quality, yielding a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views respectively. Despite being trained in a self-supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Furthermore, our analysis reveals a strong inter-task synergy between photometric scene reconstruction and semantic understanding, where semantic synthesis learning consistently enhances photometric fidelity in novel view rendering. Code will be available at https://github.com/dmucby/SPAR.

3D重建动态场景语义理解多视图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。