arXiv:2602.00866cs.AIcs.CL2026-02被引 1

用大模型预训练车载CAN数据,实现跨任务通用预测。

Foundation CAN LM: A Pretrained Language Model For Automotive CAN Data

  • 将CAN信号视为语言,统一处理离散与连续数据。
  • 单个预训练模型可适配多种汽车保险预测任务。
  • 适合自动驾驶、智能保险等需要通用表征的场景。

控制器局域网络(CAN)总线提供了丰富的车辆信号,广泛应用于碰撞检测、预测性维护和驾驶员风险建模等汽车及车险领域。然而,现有方法多针对特定任务在原始CAN数据上训练孤立模型,仅有限探索解码信号,导致表示学习难以共享,跨任务泛化能力受限。相比之下,自然语言处理与计算机视觉已因基础模型范式(大规模预训练+任务微调)实现突破。本文提出基础CAN模型,通过大规模未标注解码后的CAN信号预训练,再在多样化的车险任务中微调,实现多目标下游泛化。为支持该方法,我们设计了统一的混合离散-连续信号分词方案,并解决时序复杂性与行程特异性差异问题。实验表明,单一预训练的CAN模型能有效适应多种预测任务,验证了基础模型范式在CAN数据中的有效性,为汽车人工智能的通用表征学习开辟新方向。

原文摘要 · Abstract (English)

The Controller Area Network (CAN) bus provides a rich source of vehicular signals increasingly leveraged for applications in automotive and auto insurance domains, including collision detection, predictive maintenance, and driver risk modeling. Despite this potential, existing pipelines largely train isolated task-specific models on raw CAN data, with only limited efforts exploring decoded signals. Such fragmentation prevents shared representation learning and limits cross-task generalization. By contrast, natural language processing (NLP) and computer vision (CV) have been transformed by the foundation model paradigm: large-scale pretraining followed by task-specific adaptation. In this work, we introduce the foundation CAN model that demonstrates multi-objective downstream generalization using a single pretrained backbone. Our approach treats CAN data as a language: we pretrain on large-scale, unlabeled decoded CAN signals and fine-tune across heterogeneous auto insurance tasks. To enable this, we propose a unified tokenization scheme for mixed discrete-continuous signals and address challenges of temporal complexity and trip-specific variability. Our results show that one pretrained CAN model can adapt effectively to diverse predictive tasks, validating that the foundation modeling paradigm, proven in NLP and CV, also holds for CAN data. This establishes a new direction for generalizable representation learning in automotive AI.

汽车AI基础模型信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。