arXiv:2511.14970cs.CVcs.AI2025-11被引 2

用边缘信息融合语义与深度特征,提升透明物体的感知效果。

EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects

  • 通过边缘引导的空间注意力机制融合语义与几何特征
  • 在Syn-TODD和ClearPose上深度精度超越当前最佳方法(MODEST)
  • 无需真实深度标签,渐进式训练提升学习效率

透明物体感知仍是计算机视觉中的重大挑战,因其透明性干扰了深度估计与语义分割。现有方法虽采用多任务学习提升鲁棒性,但跨任务负向干扰常导致性能下降。本文提出边缘引导空间注意力(EGSA)融合机制,通过引入边界信息缓解语义与几何特征融合中的破坏性交互。在Syn-TODD和ClearPose两个基准上,EGSA在深度估计上持续优于当前最优方法MODEST,且在透明区域提升最为显著,同时保持优异的分割性能。此外,本文提出一种多模态渐进式训练策略:从RGB图像提取的边缘开始学习,逐步过渡到由预测深度图生成的边缘。该策略使系统先利用RGB图像中的丰富纹理进行初始化,再转向更相关的深度几何内容,且无需训练时使用真实深度标签。两项贡献共同表明,边缘引导融合是一种可有效提升透明物体感知能力的稳健方法。

原文摘要 · Abstract (English)

Transparent object perception remains a major challenge in computer vision research, as transparency confounds both depth estimation and semantic segmentation. Recent work has explored multi-task learning frameworks to improve robustness, yet negative cross-task interactions often hinder performance. In this work, we introduce Edge-Guided Spatial Attention (EGSA), a fusion mechanism designed to mitigate destructive interactions by incorporating boundary information into the fusion between semantic and geometric features. On both Syn-TODD and ClearPose benchmarks, EGSA consistently improved depth accuracy over the current state of the art method (MODEST), while preserving competitive segmentation performance, with the largest improvements appearing in transparent regions. Besides our fusion design, our second contribution is a multi-modal progressive training strategy, where learning transitions from edges derived from RGB images to edges derived from predicted depth images. This approach allows the system to bootstrap learning from the rich textures contained in RGB images, and then switch to more relevant geometric content in depth maps, while it eliminates the need for ground-truth depth at training time. Together, these contributions highlight edge-guided fusion as a robust approach capable of improving transparent object perception.

透明物体深度估计多任务学习边缘引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。