arXiv:2512.07034cs.CVcs.AI2025-12中稿 · WACV 2026被引 2

利用边界和反光特征提升玻璃等透明物体的分割精度

Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues

论文配图:Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
图 1 · 摘自论文原文
  • 设计边界与反光增强模块,协同捕捉透明物体视觉线索
  • 在多个数据集上提升性能,最高增益达13.1% mIoU
  • 适合需要高精度透明物体分割的研究与工业应用

玻璃是日常生活中常见的固体材料,但因其透明性和反射性,现有分割方法难以将其与不透明物体区分开。尽管人类感知依赖于边界和反射物体特征来识别玻璃,但现有研究尚未充分融合这两种视觉线索。为此,我们提出在金字塔式视觉变换器架构中引入边界特征增强与反射特征增强模块,以相互促进的方式建模透明物体特性。所提框架TransCues为基于编码器-解码器的金字塔变换器结构。实验表明,两个模块可有效协同,显著提升在多个基准数据集上的表现,包括玻璃对象语义分割、镜面对象语义分割及通用分割数据集。方法在Trans10K-v2上实现+4.2% mIoU,MSD上+5.6% mIoU,RGBD-Mirror上+10.1% mIoU,TROSD上+13.1% mIoU,Stanford2D3D上+8.3% mIoU,充分验证了其对玻璃类物体的有效性。

原文摘要 · Abstract (English)

Glass is a prevalent material among solid objects in everyday life, yet segmentation methods struggle to distinguish it from opaque materials due to its transparency and reflection. While it is known that human perception relies on boundary and reflective-object features to distinguish glass objects, the existing literature has not yet sufficiently captured both properties when handling transparent objects. Hence, we propose incorporating both of these powerful visual cues via the Boundary Feature Enhancement and Reflection Feature Enhancement modules in a mutually beneficial way. Our proposed framework, TransCues, is a pyramidal transformer encoder-decoder architecture to segment transparent objects. We empirically show that these two modules can be used together effectively, improving overall performance across various benchmark datasets, including glass object semantic segmentation, mirror object semantic segmentation, and generic segmentation datasets. Our method outperforms the state-of-the-art by a large margin, achieving +4.2% mIoU on Trans10K-v2, +5.6% mIoU on MSD, +10.1% mIoU on RGBD-Mirror, +13.1% mIoU on TROSD, and +8.3% mIoU on Stanford2D3D, showing the effectiveness of our method against glass objects.

透明物体分割视觉感知金字塔变换器边界增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。