arXiv:2602.02063cs.HCcs.AI2026-02ACL

用视觉语言模型自动优化自动驾驶交互动作设计。

See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action Designers

  • 用视觉语言模型提供视觉反馈,闭环改进大模型生成的交互动作。
  • 在三种交互模态上均优于固定提示和人工设计基线。
  • 无需人工标注,适合大规模自动化交互设计。

自动驾驶车辆缺乏与道路使用者的自然沟通渠道,外部人机界面(eHMI)对于传达意图和维护共享环境中的信任至关重要。然而,多数eHMI研究依赖开发者手工设计的消息-动作配对,难以适应多样动态的交通场景。一个有前景的替代方案是使用大语言模型(LLM)作为动作设计师,生成上下文相关的eHMI动作,但这类设计缺乏感知验证,通常依赖固定提示或昂贵的人工标注反馈进行优化。本文提出See2Refine,一种无需人工参与的闭环框架,利用视觉语言模型(VLM)的感知评估作为自动视觉反馈,持续优化基于LLM的eHMI动作设计师。给定驾驶上下文和候选动作,VLM评估动作的感知适当性,并将反馈用于迭代修正设计师输出,实现无监督系统性优化。我们在三种eHMI模态(lightbar、eyes、arm)和多个LLM模型规模下评估该框架。结果表明,其在多种VLM指标和真人实验中均显著优于仅使用提示的LLM基线和人工设定基线。改进效果跨模态泛化良好,且VLM评估与人类偏好高度一致,证明See2Refine在可扩展动作设计中的鲁棒性与有效性。

原文摘要 · Abstract (English)

Automated vehicles lack natural communication channels with other road users, making external Human-Machine Interfaces (eHMIs) essential for conveying intent and maintaining trust in shared environments. However, most eHMI studies rely on developer-crafted message-action pairs, which are difficult to adapt to diverse and dynamic traffic contexts. A promising alternative is to use Large Language Models (LLMs) as action designers that generate context-conditioned eHMI actions, yet such designers lack perceptual verification and typically depend on fixed prompts or costly human-annotated feedback for improvement. We present See2Refine, a human-free, closed-loop framework that uses vision-language model (VLM) perceptual evaluation as automated visual feedback to improve an LLM-based eHMI action designer. Given a driving context and a candidate eHMI action, the VLM evaluates the perceived appropriateness of the action, and this feedback is used to iteratively revise the designer's outputs, enabling systematic refinement without human supervision. We evaluate our framework across three eHMI modalities (lightbar, eyes, and arm) and multiple LLM model sizes. Across settings, our framework consistently outperforms prompt-only LLM designers and manually specified baselines in both VLM-based metrics and human-subject evaluations. Results further indicate that the improvements generalize across modalities and that VLM evaluations are well aligned with human preferences, supporting the robustness and effectiveness of See2Refine for scalable action design.

自动驾驶大模型交互设计视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。