用视觉语言模型理解汽车界面,支持跨设计自适应交互。
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
- 基于Molmo-7B和LoRa微调,结合视觉定位与推理能力。
- 在AutomotiveUI-Bench-4K上达80.8%准确率,比基线高5.6%。
- 低成本部署,适合车载系统研发与智能交互设计者。
现代车载信息娱乐系统需应对频繁的用户界面更新与多样化设计。本文提出一个视觉语言框架,实现对汽车界面的理解与交互,支持跨设计自适应。为此公开了AutomotiveUI-Bench-4K数据集,包含998张图像与4,208个标注,并提供训练数据生成流程。采用基于Molmo-7B的模型,通过低秩适配(LoRa)进行微调,构建具备评估与推理能力的评估型大动作模型(ELAM)。该模型在AutomotiveUI-Bench-4K上表现优异,平均准确率达80.8%,在ScreenSpot任务上较基线提升5.6%。其性能接近或超越桌面、移动端及网页专用模型,且仅在车载领域训练。方法成本低,可在消费级GPU上部署。
原文摘要 · Abstract (English)
Modern automotive infotainment systems necessitate intelligent and adaptive solutions to manage frequent User Interface (UI) updates and diverse design variations. This work introduces a vision-language framework to facilitate the understanding of and interaction with automotive UIs, enabling seamless adaptation across different UI designs. To support research in this field, AutomotiveUI-Bench-4K, an open-source dataset comprising 998 images with 4,208 annotations, is also released. Additionally, a data pipeline for generating training data is presented. A Molmo-7B-based model is fine-tuned using Low-Rank Adaptation (LoRa), incorporating generated reasoning along with visual grounding and evaluation capabilities. The fine-tuned Evaluative Large Action Model (ELAM) achieves strong performance on AutomotiveUI-Bench-4K (model and dataset are available on Hugging Face). The approach demonstrates strong cross-domain generalization, including a +5.6% improvement on ScreenSpot over the baseline model. An average accuracy of 80.8% is achieved on ScreenSpot, closely matching or surpassing specialized models for desktop, mobile, and web, despite being trained primarily on the automotive domain. This research investigates how data collection and subsequent fine-tuning can lead to AI-driven advancements in automotive UI understanding and interaction. The applied method is cost-efficient, and fine-tuned models can be deployed on consumer-grade GPUs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。