用大模型自动设计自动驾驶车外人机界面动作,提升沟通适应性。
Automating eHMI Action Design with LLMs for Automated Vehicle Communication
- 用大模型生成可执行的eHMI动作指令,结合3D渲染生成动作视频。
- 构建320条动作序列数据集,验证大模型生成动作接近人工水平。
- 提出自动评分方法,适合做自动驾驶交互设计的研究者参考。
自动驾驶车辆与道路使用者之间缺乏明确通信渠道,需依赖外部人机界面(eHMI)在不确定场景中有效传递信息。当前多数eHMI研究采用预设文本和手动设计动作,限制了其在动态现实场景中的部署。鉴于大语言模型(LLMs)的泛化与多用途能力,我们探索其作为自动动作设计工具的潜力。本文做出三项贡献:(1)提出集成LLM与3D渲染器的流水线,利用LLM生成可执行的eHMI动作并渲染动作片段;(2)收集包含8种意图消息、4种典型eHMI模态的共320条动作序列的用户评分数据集,验证推理增强型LLM能生成接近人类水平的动作;(3)引入两种自动评分机制——动作参考分(ARS)与视觉-语言模型(VLM),对18个LLM进行评估,发现VLM评分与人类偏好一致,但不同eHMI模态间存在差异。
原文摘要 · Abstract (English)
The absence of explicit communication channels between automated vehicles (AVs) and other road users requires the use of external Human-Machine Interfaces (eHMIs) to convey messages effectively in uncertain scenarios. Currently, most eHMI studies employ predefined text messages and manually designed actions to perform these messages, which limits the real-world deployment of eHMIs, where adaptability in dynamic scenarios is essential. Given the generalizability and versatility of large language models (LLMs), they could potentially serve as automated action designers for the message-action design task. To validate this idea, we make three contributions: (1) We propose a pipeline that integrates LLMs and 3D renderers, using LLMs as action designers to generate executable actions for controlling eHMIs and rendering action clips. (2) We collect a user-rated Action-Design Scoring dataset comprising a total of 320 action sequences for eight intended messages and four representative eHMI modalities. The dataset validates that LLMs can translate intended messages into actions close to a human level, particularly for reasoning-enabled LLMs. (3) We introduce two automated raters, Action Reference Score (ARS) and Vision-Language Models (VLMs), to benchmark 18 LLMs, finding that the VLM aligns with human preferences yet varies across eHMI modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。