无需源数据,融合多个预训练模型提升自动驾驶场景解析能力
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

- 通过构建类别原型记忆库,比较多源模型预测一致性
- 在无标注目标域数据上协同优化多个源模型,提升泛化性能
- 适合缺乏源数据但需跨环境鲁棒感知的自动驾驶系统
可靠的道路场景解析是自动驾驶车辆在开放动态环境中运行的基础能力。然而,由于像素级标注成本高,且源域数据常因隐私、安全或所有权问题无法获取,模型迁移到新部署环境仍具挑战性。现有无源域自适应方法通常依赖单一预训练源模型,易受源域偏差影响,难以应对多样道路布局、光照、天气与交通条件。本文提出一种无源域自适应协作框架(UCDA),在不访问原始源样本的前提下,将多个预训练源模型的互补知识迁移至统一目标模型。UCDA通过构建类级别原型记忆库,基于原型相似性估计不同源模型预测的可靠性,缓解模型间置信度尺度不一致问题。在此基础上,采用两阶段迁移策略:先在无标签目标域驾驶数据上通过正负一致性约束协同优化多个源模型;再将其验证后的专长知识蒸馏至可部署的目标模型。在公开驾驶场景数据集及真实自动驾驶平台采集数据上的综合评估表明,UCDA能有效整合多源互补知识,显著提升目标域场景解析的可靠性与跨环境泛化能力。
原文摘要 · Abstract (English)
Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting perception models to new deployment domains remains challenging because pixel-level annotations are expensive to obtain, while source-domain data are often inaccessible due to privacy, security, or ownership constraints. Existing source-free unsupervised domain adaptation methods typically rely on a single pre-trained source model, which makes the adapted perception system vulnerable to source-specific biases and limits its robustness under diverse road layouts, illumination conditions, weather patterns, and traffic conditions. This article presents an unsupervised collaborative domain adaptation (UCDA) framework for driving scene parsing in a source-free setting, which transfers complementary knowledge from multiple pre-trained source models to a unified target model without accessing any original source samples. To compare predictions from independently trained models, UCDA constructs a class-level prototype memory bank and estimates cross-model prediction reliability through prototype similarity, reducing the effect of inconsistent confidence scales across source models. Based on the resulting complementary supervision, UCDA adopts a two-stage transfer strategy: multiple source models are first refined on unlabeled target-domain driving data through collaborative optimization with positive and negative consistency constraints, and their validated expertise is then distilled into a single deployable target model. Comprehensive evaluations on public driving-scene datasets and real-world data collected from an autonomous vehicle platform demonstrate that UCDA effectively consolidates complementary multi-source knowledge, improving target-domain scene parsing reliability and generalization across diverse driving environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。