arXiv:2607.04500cs.CV2026-07

地理多样性比数据量更能提升驾驶模型跨域泛化能力

Geographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World Models

  • 用多城市数据训练世界模型,提升跨城泛化性能
  • 地理多样数据使惊喜分数降低16.5%(0.228对0.273)
  • 适合关注自动驾驶模型鲁棒性的研究者

自监督潜在世界模型可在无标签情况下为驾驶场景分配惊奇度评分。本文通过受控迁移实验探讨:在某一地理区域训练的模型,能否将复杂性认知推广至未见城市与传感器配置?我们以nuPlan数据(匹兹堡、波士顿、新加坡)训练基于JEPA的世界模型,并在迈阿密和奥斯汀的Argoverse 2验证集上进行零样本评估。结果表明,使用地理多样数据训练的模型显著优于仅使用单一地理数据的模型。在每组63,000个场景(各3个种子)的对比实验中,综合训练使平均惊奇度降低16.5%(0.228 ± 0.015 对比 0.273 ± 0.008)。值得注意的是,仅使用20万条来自单一地理的AV2数据(是综合数据的3倍),其惊奇度仍高达0.264,高于综合训练模型,表明地理多样性比数据总量更关键。

原文摘要 · Abstract (English)

Self-supervised latent world models can assign a surprise score to driving scenarios without any human labels. A natural follow-up question is whether such a model, trained on driving data from one geographic region, can generalize its notion of complexity to unseen cities and sensor configurations. We study this question through a controlled transfer experiment: we train JEPA-based world models on nuPlan data (Pittsburgh, Boston, Singapore) and evaluate zero-shot on held-out Argoverse 2 validation scenarios from Miami and Austin. We find that models trained on geographically diverse data generalize significantly better than models trained on equal amounts of single-geography data. In a matched-scale ablation at 63,000 scenarios per condition (n=3 seeds each), combined training reduces mean surprise score by 16.5% relative to nuPlan-only training (0.228 +/- 0.015 vs 0.273 +/- 0.008). Notably, training on 200,000 AV2-only scenarios (3x more data from one geography) still produces higher surprise (0.264) than the combined 63K model, suggesting that geographic diversity is a stronger predictor of cross-domain generalization than raw data volume.

自动驾驶自监督学习跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。