arXiv:2606.22706cs.AI2026-06

提出新评估指标DSAIS,精准衡量车载语音干预消息的安全性与适配性。

Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion

论文配图:Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion
图 1 · 摘自论文原文
  • 融合多任务识别输出与风险融合机制,动态生成干预提示。
  • 在AIDE数据集上,评分一致性达ICC 0.798-0.840,显著优于基线9.1%。
  • 发现驾驶员情绪识别最关键,7B-9B小模型性能超API大模型,适合车载部署。

现有驾驶员干预系统依赖听觉警报和固定模板,未能充分利用多任务识别输出。通用评价指标如BLEU和BERTScore无法捕捉干预质量中的风险-紧急度匹配、认知负荷和驾驶者接受度等关键维度。本文提出驾驶员安全感知干预评分(DSAIS),通过轻量规则计算与LLM判别结合的混合架构,评估五个维度。构建端到端框架,整合四任务识别输出,通过风险融合、状态历史管理与动态提示构造实现。在AIDE数据集上,五种模型七种条件下实验表明,DSAIS在三位不同架构判别器间达到ICC 0.798–0.840,所有对照条件下Cohen's d > 1.5。多维子评分分析揭示规则基系统与LLM系统间存在上下文适应性差距,多任务融合使上下文相关性提升9.1%。消融实验证明各组件均贡献于上下文相关性,子评分分解显示增益可叠加。驾驶员情绪识别为最关键上游因素,7B–9B参数的紧凑本地LLM性能优于API模型,为车载部署提供实用指导。

原文摘要 · Abstract (English)

Existing driver intervention systems rely on auditory alerts and fixed templates, failing to leverage multi-task recognition outputs. General-purpose metrics such as BLEU and BERTScore cannot capture intervention-specific quality dimensions including risk-urgency alignment, cognitive load, and driver acceptability. In this paper, we propose the Driver Safety-Aware Intervention Score (DSAIS), a domain-specific metric evaluating five dimensions through a hybrid architecture combining lightweight rule-based computation with LLM Judge evaluation, together with an end-to-end framework integrating four-task recognition outputs into an LLM through risk fusion, state history management, and dynamic prompt construction. Experiments on the AIDE dataset with five models and seven conditions demonstrate that DSAIS achieves ICC 0.798-0.840 across three architecturally distinct judges and Cohen's d > 1.5 across all control conditions. Multi-dimensional sub-score analysis quantifies the contextual adaptability gap between rule-based and LLM-based systems, revealing that multi-task integration improves contextual relevance by 9.1% over rule-based baselines. Ablation experiments demonstrate that each framework component contributes to contextual relevance, with sub-score decomposition revealing gains that aggregate scoring masks. Driver emotion recognition is identified as the most critical upstream factor, and compact local LLMs (7B--9B parameters) achieve quality superior to API-based models, providing practical design guidelines for in-vehicle deployment.

LLM评估智能驾驶多任务融合车载系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。