用闲置自动驾驶算力帮外卖员精准找停车位,省时又赚钱。
ParkSense: Where Should a Delivery Driver Park? Leveraging Idle AV Compute and Vision-Language Models
- 利用自动驾驶车辆空闲算力,运行视觉语言模型分析影像找入口和合法停车位。
- 70亿参数量化模型在普通硬件上4-8秒完成推理,显著提升停车效率。
- 适合关注自动驾驶资源复用、最后一公里物流优化的研究者与从业者。
找停车位占用了外卖配送的不成比例时间,但目前尚无系统能针对商家出入口精准选择停车点。我们提出 ParkSense 框架,利用自动驾驶在低风险状态(如红灯排队、交通拥堵、停车场慢行)下的闲置计算资源,对预先缓存的卫星图与街景图像运行视觉语言模型(VLM),识别商家入口和合法停车区域。我们形式化了面向配送的高精度停车问题(DAPP),表明一个量化后的70亿参数VLM可在HW4级硬件上于4-8秒内完成推理,并估算美国每位司机年收入可增加3,000至8,000美元。文中还提出了五个开放研究方向,探索自动驾驶、计算机视觉与最后一公里物流交叉领域的未开发潜力。
原文摘要 · Abstract (English)
Finding parking consumes a disproportionate share of food delivery time, yet no system addresses precise parking-spot selection relative to merchant entrances. We propose ParkSense, a framework that repurposes idle compute during low-risk AV states -- queuing at red lights, traffic congestion, parking-lot crawl -- to run a Vision-Language Model (VLM) on pre-cached satellite and street view imagery, identifying entrances and legal parking zones. We formalize the Delivery-Aware Precision Parking (DAPP) problem, show that a quantized 7B VLM completes inference in 4-8 seconds on HW4-class hardware, and estimate annual per-driver income gains of 3,000-8,000 USD in the U.S. Five open research directions are identified at this unexplored intersection of autonomous driving, computer vision, and last-mile logistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。