用深度学习生成更智能的导航语音,减少走错路。
Revolutionizing Turn-by-Turn Navigation with Cloud-Edge Deep Learning

- 将导航指令拆解为模块化元素,用Transformer+MoE联合建模
- 实测偏离路线比例显著降低,导航更精准有效
- 云边协同架构支持实时应用,适合智能驾驶系统
转弯导航系统是现代驾驶体验的核心,提供实时语音指引以确保安全抵达目的地。然而,现有语音指令多依赖规则方法,在信息量与认知负荷之间难以平衡,复杂环境下易导致驾驶员困惑或错过转向。为此,我们首次将导航指令生成建模为多任务学习问题,将音频内容分解为模块化组合;提出一种新型深度学习框架,融合Transformer强大的时空处理能力与混合专家(MoE)的多任务学习优势,生成实时、上下文感知的语音指令。采用云边协同架构应对模型计算需求,保障可扩展性与实时性能。真实世界实验表明,该方法显著降低车辆偏离路线的比例,提供更清晰有效的语音指引。这是深度学习在驾驶语音导航中的首次大规模应用,标志着智能交通与驾驶辅助技术的重大进步。
原文摘要 · Abstract (English)
Turn-by-turn (TBT) navigation systems are integral to modern driving experiences, providing real-time audio instructions to guide drivers safely to destinations. However, existing audio instruction policy often relies on rule-based approaches that struggle to balance informational content with cognitive load, potentially leading to driver confusion or missed turns in complex environments. To overcome these difficulties, we first model the generation of navigation instructions as a multi-task learning problem by decomposing the audio content into combinations of modular elements. Then, we propose a novel deep learning framework that leverages the powerful spatiotemporal information processing capabilities of Transformers and the strong multi-task learning abilities of Mixture of Experts (MoE) to generate real-time, context-aware audio instructions for TBT driving navigation. A cloud-edge collaborative architecture is implemented to handle the computational demands of the model, ensuring scalability and real-time performance for practical applications. Experimental results in the real world demonstrate that the proposed method significantly reduces the yaw rate (the proportion of vehicles deviating from navigation routes) compared to traditional methods, delivering clearer and more effective audio instructions. This is the first large-scale application of deep learning in driving audio navigation, marking a substantial advancement in intelligent transportation and driving assistance technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。