arXiv:2603.02648cs.CV2026-03中稿 · ISCAS 2026

通过频域增强与多尺度精修,提升透明物体分割精度

SEP-YOLO: Fourier-Domain Feature Representation for Transparent Object Instance Segmentation

  • 在频域分离并增强微弱的高频边界特征
  • 多尺度精修流实现深层语义特征的精准对齐与定位
  • 首个针对透明物体的高质量实例级标注数据集

透明物体实例分割在计算机视觉中面临严峻挑战,因其固有特性如边界模糊、对比度低以及对背景依赖性强。现有方法常因依赖强烈外观线索和清晰边界而失效。为此,我们提出SEP-YOLO,一种融合双域协同机制的新型框架。该方法引入频域细节增强模块,通过可学习复权重分离并增强微弱的高频边界成分;进一步设计多尺度空间精修流,包含内容感知对齐颈和多尺度门控精修块,以确保深层语义特征中的精确特征对齐与边界定位。同时,我们为Trans10K数据集提供了高质量实例级标注,填补了透明物体实例分割的关键数据空白。在Trans10K和GVD数据集上的大量实验表明,SEP-YOLO达到当前最优(SOTA)性能。

原文摘要 · Abstract (English)

Transparent object instance segmentation presents significant challenges in computer vision, due to the inherent properties of transparent objects, including boundary blur, low contrast, and high dependence on background context. Existing methods often fail as they depend on strong appearance cues and clear boundaries. To address these limitations, we propose SEP-YOLO, a novel framework that integrates a dual-domain collaborative mechanism for transparent object instance segmentation. Our method incorporates a Frequency Domain Detail Enhancement Module, which separates and enhances weak highfrequency boundary components via learnable complex weights. We further design a multi-scale spatial refinement stream, which consists of a Content-Aware Alignment Neck and a Multi-scale Gated Refinement Block, to ensure precise feature alignment and boundary localization in deep semantic features. We also provide high-quality instance-level annotations for the Trans10K dataset, filling the critical data gap in transparent object instance segmentation. Extensive experiments on the Trans10K and GVD datasets show that SEP-YOLO achieves state-of-the-art (SOTA) performance.

透明物体实例分割频域增强多尺度精修

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。