评测大模型在真实不确定场景下的可靠性与自我认知能力
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
- 构建车内助手场景的多轮工具调用基准,模拟用户请求模糊与信息缺失
- 前沿模型在澄清任务中一致性通过率不足50%,常提前行动或编造信息
- 突出模型对自身能力边界的认知缺陷,适合关注实际部署可靠性的研究者
现有大模型代理评估基准多聚焦理想化环境下的任务完成度,忽视真实应用场景中的可靠性。在车载语音助手等场景中,用户常提出不完整或模糊的请求,引发固有不确定性,代理需通过对话、工具调用和策略遵守来应对。我们提出CAR-bench,一个用于评估多轮、工具使用型大模型代理在车载助手领域中的一致性、不确定性处理及能力自知能力的基准。该环境包含由大模型模拟的用户、领域政策以及覆盖导航、生产力、充电和车辆控制的58个相互关联的工具。除标准任务完成外,CAR-bench引入幻觉任务(测试在缺少工具或信息时的极限认知)与歧义消解任务(要求通过澄清或内部信息获取解决不确定性)。基线结果揭示各类任务中偶发成功与持续成功之间存在巨大差距。即使前沿推理模型在歧义消解任务中一致性通过率也低于50%,因过早行动;且在幻觉任务中频繁违反策略或虚构信息以满足用户需求,凸显真实场景下更可靠、自知的模型亟待发展。
原文摘要 · Abstract (English)
Existing benchmarks for Large Language Model (LLM) agents focus on task completion under idealistic settings but overlook reliability in real-world, user-facing applications. In domains, such as in-car voice assistants, users often issue incomplete or ambiguous requests, creating intrinsic uncertainty that agents must manage through dialogue, tool use, and policy adherence. We introduce CAR-bench, a benchmark for evaluating consistency, uncertainty handling, and capability awareness in multi-turn, tool-using LLM agents in an in-car assistant domain. The environment features an LLM-simulated user, domain policies, and 58 interconnected tools spanning navigation, productivity, charging, and vehicle control. Beyond standard task completion, CAR-bench introduces Hallucination tasks that test agents' limit-awareness under missing tools or information, and Disambiguation tasks that require resolving uncertainty through clarification or internal information gathering. Baseline results reveal large gaps between occasional and consistent success on all task types. Even frontier reasoning LLMs achieve less than 50% consistent pass rate on Disambiguation tasks due to premature actions, and frequently violate policies or fabricate information to satisfy user requests in Hallucination tasks, underscoring the need for more reliable and self-aware LLM agents in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。