arXiv:2601.22529cs.CV2026-01

将分割结构融入密集预测,提升3D感知的准确性与可解释性。

SHED Light on Segmentation for Dense Prediction

  • 通过双向分层推理显式引入几何先验
  • 深度边界更清晰,语义一致性更强,跨域泛化好
  • 适合需要结构理解的机器人与三维重建任务

密集预测从单张图像推断像素级值,是3D感知与机器人技术的基础。尽管真实场景具有强结构,现有方法仍将其视为独立像素预测,常导致结构不一致。我们提出SHED,一种新型编码器-解码器架构,通过将分割融入密集预测,显式引入几何先验。通过双向分层推理,分割标记在编码器中分层聚合,在解码器中反向展开。模型仅在最终输出监督,使分割层次结构自发形成,无需显式分割标注。SHED提升了深度边界锐度与分割连贯性,表现出从合成到真实环境的强跨域泛化能力。其层次感知解码器更准确捕捉全局3D场景布局,显著改善语义分割性能。此外,该方法提升了3D重建质量,并揭示了传统像素级方法常忽略的部件级结构。

原文摘要 · Abstract (English)

Dense prediction infers per-pixel values from a single image and is fundamental to 3D perception and robotics. Although real-world scenes exhibit strong structure, existing methods treat it as an independent pixel-wise prediction, often resulting in structural inconsistencies. We propose SHED, a novel encoder-decoder architecture that enforces geometric prior explicitly by incorporating segmentation into dense prediction. By bidirectional hierarchical reasoning, segment tokens are hierarchically pooled in the encoder and unpooled in the decoder to reverse the hierarchy. The model is supervised only at the final output, allowing the segment hierarchy to emerge without explicit segmentation supervision. SHED improves depth boundary sharpness and segment coherence, while demonstrating strong cross-domain generalization from synthetic to the real-world environments. Its hierarchy-aware decoder better captures global 3D scene layouts, leading to improved semantic segmentation performance. Moreover, SHED enhances 3D reconstruction quality and reveals interpretable part-level structures that are often missed by conventional pixel-wise methods.

密集预测3D感知分割融合结构先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。