通过结构化原型正则化,提升合成数据训练的自动驾驶场景分割在真实场景的表现。
Structured prototype regularization for synthetic-to-real driving scene parsing
- 用类别专属原型强化类间分离与类内紧凑,保持特征语义结构。
- 在Cityscapes和BDD100K上平均性能超越现有方法,达78.6%和82.3%的mIoU。
- 适合做自动驾驶视觉理解、无监督域适应的研究者参考。
驾驶场景解析对自动驾驶车辆在复杂真实交通环境中的可靠运行至关重要。为降低对昂贵像素级标注的依赖,带有自动标签的合成数据集已成为主流替代方案。然而,仅在合成数据上训练的模型在真实场景中表现不佳,主要源于合成到真实的域差异。尽管无监督域适应在缩小该差距方面取得进展,但多数方法仅关注全局特征对齐,忽视了特征空间中的语义结构,导致类别间语义关系建模不足,限制了泛化能力。为此,本文提出一种新型无监督域适应框架,通过显式正则化语义特征结构,显著提升真实场景下的驾驶场景解析性能。具体而言,利用类别专属原型强制实现类间分离与类内紧凑,增强特征簇的判别性与结构一致性;采用基于熵的噪声过滤策略提高伪标签可靠性;引入像素级注意力机制进一步优化特征对齐。在代表性基准测试上的大量实验表明,所提方法持续优于近期先进方法。结果凸显了在合成到真实适应任务中保持语义结构的重要性。
原文摘要 · Abstract (English)
Driving scene parsing is critical for autonomous vehicles to operate reliably in complex real-world traffic environments. To reduce the reliance on costly pixel-level annotations, synthetic datasets with automatically generated labels have become a popular alternative. However, models trained on synthetic data often perform poorly when applied to real-world scenes due to the synthetic-to-real domain gap. Despite the success of unsupervised domain adaptation in narrowing this gap, most existing methods mainly focus on global feature alignment while overlooking the semantic structure of the feature space. As a result, semantic relations among classes are insufficiently modeled, limiting the model's ability to generalize. To address these challenges, this study introduces a novel unsupervised domain adaptation framework that explicitly regularizes semantic feature structures to significantly enhance driving scene parsing performance in real-world scenarios. Specifically, the proposed method enforces inter-class separation and intra-class compactness by leveraging class-specific prototypes, thereby enhancing the discriminability and structural coherence of feature clusters. An entropy-based noise filtering strategy improves the reliability of pseudo labels, while a pixel-level attention mechanism further refines feature alignment. Extensive experiments on representative benchmarks demonstrate that the proposed method consistently outperforms recent state-of-the-art methods. These results underscore the importance of preserving semantic structure for robust synthetic-to-real adaptation in driving scene parsing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。