arXiv:2508.06546cs.CVeess.IV2025-08ICCV被引 6

用统计置信度重评分提升多视角图像的3D场景图生成鲁棒性

Statistical Confidence Rescoring for Robust 3D Scene Graph Generation from Multi-View Images

  • 基于多视角图像,通过语义掩码过滤背景噪声
  • 引入邻域统计先验优化节点与边的预测置信度
  • 适用于无真实3D标注的复杂场景图生成任务

当前3D语义场景图估计方法依赖真实3D标注来准确预测目标物体、谓词和关系。在缺乏3D真值的情况下,我们探索仅使用多视角RGB图像完成该任务。为获得稳健特征以实现准确的场景图估计,需克服由预测深度图重构的伪点云几何带来的噪声,并减少多视角图像特征中的背景噪声。关键在于通过邻域关系丰富节点与边的语义和空间信息。我们利用语义掩码引导特征聚合以过滤背景特征,并设计新方法融合邻近节点信息以增强场景图估计的鲁棒性。此外,我们利用训练集统计量计算的显式统计先验,根据一跳邻域信息对节点和边的预测进行精炼。实验表明,该方法在仅使用多视角图像作为输入时优于现有方法。

原文摘要 · Abstract (English)

Modern 3D semantic scene graph estimation methods utilize ground truth 3D annotations to accurately predict target objects, predicates, and relationships. In the absence of given 3D ground truth representations, we explore leveraging only multi-view RGB images to tackle this task. To attain robust features for accurate scene graph estimation, we must overcome the noisy reconstructed pseudo point-based geometry from predicted depth maps and reduce the amount of background noise present in multi-view image features. The key is to enrich node and edge features with accurate semantic and spatial information and through neighboring relations. We obtain semantic masks to guide feature aggregation to filter background features and design a novel method to incorporate neighboring node information to aid robustness of our scene graph estimates. Furthermore, we leverage on explicit statistical priors calculated from the training summary statistics to refine node and edge predictions based on their one-hop neighborhood. Our experiments show that our method outperforms current methods purely using multi-view images as the initial input. Our project page is available at https://qixun1.github.io/projects/SCRSSG.

3D场景图多视角图像鲁棒性统计先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。