语音交互的智能车载助手,驾驶时认知负担与免提通话相当。
Visual and Cognitive Demands of a Large Language Model-Powered In-vehicle Conversational Agent
- 用语音交互的LLM助手进行对话,认知负荷与免提通话接近。
- 多轮对话中认知负荷保持稳定,眼动数据显示视觉注意力始终在安全范围。
- 适合关注车载AI安全性、人机交互设计的研究者和开发者。
驾驶员分心仍是机动车事故的主要原因,亟需对新型车载技术进行严格评估。本研究评估了先进大语言模型(LLM)对话代理Gemini Live在实际道路驾驶中的视觉与认知需求,对比免提电话通话、低负载基准的可视化导航(转弯指引)以及高负载锚点任务操作跨度测试(OSPAN)。32名持证驾驶员完成了五项次级任务,通过检测反应任务(DRT)测量认知负荷,眼动追踪分析视觉注意力,主观工作量评分辅助验证。结果显示,无论是单轮还是多轮交互,Gemini Live与免提电话的认知负荷水平相似,介于转弯指引与OSPAN之间。探索性分析表明,长时间多轮对话中认知负荷保持稳定。所有任务的平均注视时间均远低于2秒的安全阈值,确认视觉需求较低。此外,驾驶员在完成任务期间,会将更长的注视时间用于道路,尤其是在语音交互时,尽管总离眼时间较长,但实际影响较小。主观评价与客观数据一致,参与者普遍认为使用Gemini Live时努力程度、任务需求和感知分心度均较低。研究证明,通过语音接口实现的先进LLM对话代理,在驾驶环境中带来的认知与视觉负担与已知低风险免提基准相当,支持其安全部署。
原文摘要 · Abstract (English)
Driver distraction remains a leading contributor to motor vehicle crashes, necessitating rigorous evaluation of new in-vehicle technologies. This study assessed the visual and cognitive demands associated with an advanced Large Language Model (LLM) conversational agent (Gemini Live) during on-road driving, comparing it against handsfree phone calls, visual turn-by-turn guidance (low load baseline), and the Operation Span (OSPAN) task (high load anchor). Thirty-two licensed drivers completed five secondary tasks while visual and cognitive demands were measured using the Detection Response Task (DRT) for cognitive load, eye-tracking for visual attention, and subjective workload ratings. Results indicated that Gemini Live interactions (both single-turn and multi-turn) and hands-free phone calls shared similar levels of cognitive load, between that of visual turn-by-turn guidance and OSPAN. Exploratory analysis showed that cognitive load remained stable across extended multi-turn conversations. All tasks maintained mean glance durations well below the well-established 2-second safety threshold, confirming low visual demand. Furthermore, drivers consistently dedicated longer glances to the roadway between brief off-road glances toward the device during task completion, particularly during voice-based interactions, rendering longer total-eyes-off-road time findings less consequential. Subjective ratings mirrored objective data, with participants reporting low effort, demands, and perceived distraction for Gemini Live. These findings demonstrate that advanced LLM conversational agents, when implemented via voice interfaces, impose cognitive and visual demands comparable to established, low-risk hands-free benchmarks, supporting their safe deployment in the driving environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。