首个面向车载界面的多模态智能体评测基准,助力安全交互研究。
Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- 构建185个参数化任务,覆盖控车、意图理解与安全场景。
- 引入地理信息增强模型,在安全任务上成功率显著提升。
- 适合自动驾驶、人机交互领域研究者参考使用。
多模态智能体在通用图形界面中表现优异,但在车载系统中应用仍不充分。车内图形界面面临注意力有限、安全要求严格及位置相关交互复杂等挑战。为此,我们提出Automotive-ENV,首个高保真车载界面评测基准与交互环境。该平台定义了185个参数化任务,涵盖显式控制、隐式意图理解与安全感知任务,并提供结构化多模态观测与精确程序化验证,支持可复现评估。基于此基准,我们提出ASURADA,一种融合地理信息的多模态智能体,利用GPS上下文动态调整行为,适应位置、环境条件与区域驾驶规范。实验表明,地理信息显著提升安全任务成功率,凸显位置上下文在车载环境中的关键作用。我们将开源Automotive-ENV及其全部任务与评测工具,推动安全自适应车载智能体的发展。
原文摘要 · Abstract (English)
Multimodal agents have demonstrated strong performance in general GUI interactions, but their application in automotive systems has been largely unexplored. In-vehicle GUIs present distinct challenges: drivers' limited attention, strict safety requirements, and complex location-based interaction patterns. To address these challenges, we introduce Automotive-ENV, the first high-fidelity benchmark and interaction environment tailored for vehicle GUIs. This platform defines 185 parameterized tasks spanning explicit control, implicit intent understanding, and safety-aware tasks, and provides structured multimodal observations with precise programmatic checks for reproducible evaluation. Building on this benchmark, we propose ASURADA, a geo-aware multimodal agent that integrates GPS-informed context to dynamically adjust actions based on location, environmental conditions, and regional driving norms. Experiments show that geo-aware information significantly improves success on safety-aware tasks, highlighting the importance of location-based context in automotive environments. We will release Automotive-ENV, complete with all tasks and benchmarking tools, to further the development of safe and adaptive in-vehicle agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。