用生成式AI让车载系统能听懂场景和对话,实时响应司机指令。
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
- 融合视觉理解与大模型,实现基于场景的自然语言交互。
- 支持多轮对话并确认司机意图,生成结构化控制指令。
- 无需微调模型,适合需要灵活人机协作的智能驾驶场景。
尽管自动驾驶技术持续进步,当前高级驾驶辅助系统(ADAS)仍难以理解环境上下文或通过自然语言与驾驶员互动。这些系统通常依赖预设逻辑,缺乏对话式交互能力,在动态环境或需适应驾驶员意图时表现僵硬。本文提出场景感知对话式ADAS(SC-ADAS),一个模块化框架,集成大语言模型、视觉转文本解析及结构化函数调用,实现实时、可解释且自适应的驾驶辅助。SC-ADAS支持基于视觉与传感器上下文的多轮对话,可自然语言提供建议并由驾驶员确认执行ADAS控制。在CARLA仿真器中部署,结合云端生成式AI,系统在不进行模型微调的情况下,将确认的用户意图转化为结构化指令。评估涵盖场景感知、对话能力和多轮回溯交互,结果显示视觉上下文检索带来延迟增加,对话历史累积导致令牌数量增长。这些结果证明了将对话推理、场景感知与模块化控制结合,可为下一代智能驾驶辅助提供可行性方案。
原文摘要 · Abstract (English)
While autonomous driving technologies continue to advance, current Advanced Driver Assistance Systems (ADAS) remain limited in their ability to interpret scene context or engage with drivers through natural language. These systems typically rely on predefined logic and lack support for dialogue-based interaction, making them inflexible in dynamic environments or when adapting to driver intent. This paper presents Scene-Aware Conversational ADAS (SC-ADAS), a modular framework that integrates Generative AI components including large language models, vision-to-text interpretation, and structured function calling to enable real-time, interpretable, and adaptive driver assistance. SC-ADAS supports multi-turn dialogue grounded in visual and sensor context, allowing natural language recommendations and driver-confirmed ADAS control. Implemented in the CARLA simulator with cloud-based Generative AI, the system executes confirmed user intents as structured ADAS commands without requiring model fine-tuning. We evaluate SC-ADAS across scene-aware, conversational, and revisited multi-turn interactions, highlighting trade-offs such as increased latency from vision-based context retrieval and token growth from accumulated dialogue history. These results demonstrate the feasibility of combining conversational reasoning, scene perception, and modular ADAS control to support the next generation of intelligent driver assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。