用合成数据混合提升城市道路图像的实时泛化能力
SynthGenNet: a self-supervised approach for test-time generalization using synthetic multi-source domain mixing of street view images
- 通过合成多源图像混合增强模型适应性
- 在真实数据集上达到50% mIoU,超越单源方法
- 无需目标域标签,适合复杂城市场景应用
非结构化城市环境因布局复杂多样,给场景理解与泛化带来挑战。我们提出SynthGenNet,一种基于合成多源图像混合的自监督师生架构,实现鲁棒的测试时领域泛化。核心贡献包括新型ClassMix++算法,可在保持语义一致性的前提下融合多个合成源的标注数据;引入基于源真值的地面掩码一致性损失(GMC),提升跨域预测一致性与特征对齐;在学生网络中集成伪标签引导对比学习(PLGCL)机制,通过教师网络迭代知识蒸馏实现领域不变特征学习。该自监督策略显著提升预测准确率,缓解现实世界变化影响,缩小模拟到真实域差距,并减少对目标域标注数据依赖,即使在复杂城市区域也表现优异。实验显示,模型在印度驾驶数据集(IDD)等真实数据集上取得50%的平均交并比(mIoU),超越现有基于单一来源的方法。
原文摘要 · Abstract (English)
Unstructured urban environments present unique challenges for scene understanding and generalization due to their complex and diverse layouts. We introduce SynthGenNet, a self-supervised student-teacher architecture designed to enable robust test-time domain generalization using synthetic multi-source imagery. Our contributions include the novel ClassMix++ algorithm, which blends labeled data from various synthetic sources while maintaining semantic integrity, enhancing model adaptability. We further employ Grounded Mask Consistency Loss (GMC), which leverages source ground truth to improve cross-domain prediction consistency and feature alignment. The Pseudo-Label Guided Contrastive Learning (PLGCL) mechanism is integrated into the student network to facilitate domain-invariant feature learning through iterative knowledge distillation from the teacher network. This self-supervised strategy improves prediction accuracy, addresses real-world variability, bridges the sim-to-real domain gap, and reliance on labeled target data, even in complex urban areas. Outcomes show our model outperforms the state-of-the-art (relying on single source) by achieving 50% Mean Intersection-Over-Union (mIoU) value on real-world datasets like Indian Driving Dataset (IDD).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。