arXiv:2508.19651cs.CV2025-08

用云端协同方案让车载系统高效识别车内物品,性能远超大模型。

Scalable Object Detection in the Car Interior With Vision Foundation Models

  • 分步计算:车载端与云端协作,用轻量模型+微调提升效率。
  • 微调后模型准确率达89%,比基础版提升71%,胜过GPT-4o近20%。
  • 减少幻觉三倍,适合资源受限的智能汽车场景部署。

车舱内智能任务如识别和定位外部引入物体对个人助手响应质量至关重要。然而车载系统算力高度受限,难以直接部署此类方案。为此,我们提出新型车载物体检测与定位(ODAL)框架,利用视觉基础模型通过分布式架构,将计算任务拆分至车载端与云端,突破车载资源限制。为评估模型性能,我们引入ODALbench新指标,全面衡量检测与定位能力。对比分析显示,采用微调的轻量级LLaVA 1.5 7B模型在该框架下表现优异:其ODAL$_{score}$达89%,较基线提升71%,并优于GPT-4o近20%;同时大幅降低幻觉问题,实现三倍于GPT-4o的ODAL$_{SNR}$,展现出显著优势。

原文摘要 · Abstract (English)

AI tasks in the car interior like identifying and localizing externally introduced objects is crucial for response quality of personal assistants. However, computational resources of on-board systems remain highly constrained, restricting the deployment of such solutions directly within the vehicle. To address this limitation, we propose the novel Object Detection and Localization (ODAL) framework for interior scene understanding. Our approach leverages vision foundation models through a distributed architecture, splitting computational tasks between on-board and cloud. This design overcomes the resource constraints of running foundation models directly in the car. To benchmark model performance, we introduce ODALbench, a new metric for comprehensive assessment of detection and localization.Our analysis demonstrates the framework's potential to establish new standards in this domain. We compare the state-of-the-art GPT-4o vision foundation model with the lightweight LLaVA 1.5 7B model and explore how fine-tuning enhances the lightweight models performance. Remarkably, our fine-tuned ODAL-LLaVA model achieves an ODAL$_{score}$ of 89%, representing a 71% improvement over its baseline performance and outperforming GPT-4o by nearly 20%. Furthermore, the fine-tuned model maintains high detection accuracy while significantly reducing hallucinations, achieving an ODAL$_{SNR}$ three times higher than GPT-4o.

目标检测车载AI基础模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。