用无标签数据提升自动驾驶语言模型,5%标注数据即可逼近全量训练效果
Unlock the Power of Unlabeled Data in Language Driving Model
- 通过模板生成问题,为无标签数据构造伪答案
- 自洽性优化提升伪标注质量,实现半监督训练
- 仅需5%标注数据即达44.85%,融合无标签数据后达54.27%
近期基于视觉的大语言模型在自动驾驶领域发展迅速,但其性能高度依赖大规模高质量标注数据,成本高昂。为此,我们提出一种半监督学习方法,挖掘海量无标签数据的价值以提升语言驱动模型。首先,设计一系列基于模板的提示,从无标签数据中提取场景信息,生成问题并基于少量标注数据训练的模型生成伪答案;其次,提出自洽性精炼方法,提升伪标注质量,用于后续模型训练。基于预训练的VisionLLM(如InternVL),构建强大的语言驾驶模型(LDM),在DriveLM基准上表现优于现有最先进方法。大量实验表明,本方法仅使用5%标注数据即可取得良好效果,在引入无标签数据后性能提升至54.27%,接近全数据训练模型(60.68%)的表现。
原文摘要 · Abstract (English)
Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quality annotated data, which is costly and labor-intensive. To address this issue, we propose unlocking the value of abundant yet unlabeled data to improve the language-driving model in a semi-supervised learning manner. Specifically, we first introduce a series of template-based prompts to extract scene information, generating questions that create pseudo-answers for the unlabeled data based on a model trained with limited labeled data. Next, we propose a Self-Consistency Refinement method to improve the quality of these pseudo-annotations, which are later used for further training. By utilizing a pre-trained VisionLLM (e.g., InternVL), we build a strong Language Driving Model (LDM) for driving scene question-answering, outperforming previous state-of-the-art methods. Extensive experiments on the DriveLM benchmark show that our approach performs well with just 5% labeled data, achieving competitive performance against models trained with full datasets. In particular, our LDM achieves 44.85% performance with limited labeled data, increasing to 54.27% when using unlabeled data, while models trained with full datasets reach 60.68% on the DriveLM benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。