评测大模型汽车助手在手册查询中的失效问题,发现关键漏洞。
DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant

- 用大模型生成测试用例,检测汽车助手遗漏警告的场景。
- 最佳工具发现17种不同失效模式,覆盖多类用户提问。
- 适合自动驾驶与智能座舱安全验证的研究者参考。
本报告总结了首届大型语言模型(LLM)测试竞赛的结果,该竞赛作为ICSE 2026年DeepTest研讨会的一部分举行。四个工具参与评测一个基于LLM的汽车手册信息检索应用,目标是识别系统未能正确提及手册中警告信息的用户输入。测试方案根据暴露故障的有效性及发现故障测试用例的多样性进行评估。报告详述了实验方法、参赛工具及其结果。
原文摘要 · Abstract (English)
This report summarizes the results of the first edition of the Large Language Model (LLM) Testing competition, held as part of the DeepTest workshop at ICSE 2026. Four tools competed in benchmarking an LLM-based car manual information retrieval application, with the objective of identifying user inputs for which the system fails to appropriately mention warnings contained in the manual. The testing solutions were evaluated based on their effectiveness in exposing failures and the diversity of the discovered failure-revealing tests. We report on the experimental methodology, the competitors, and the results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。