对比大模型与传统规划器的人类信任度,发现正确性最关键。
Evaluating Human Trust in LLM-Based Planners: A Preliminary Study
- 通过用户实验比较大模型与传统规划器的信任表现
- 正确性显著影响信任,解释提升准确率但不明显增加信任
- 计划迭代优化或可增强信任,适合人机协作研究者
大型语言模型(LLMs)在规划任务中的应用日益广泛,具备生成解释和迭代优化等传统规划器所不具备的能力。然而,在基于大模型的规划场景中,信任这一关键因素仍缺乏深入研究。本研究通过在规划领域定义语言(PDDL)环境中开展用户实验,对比了人类对大模型规划器与传统规划器的信任水平。结合主观测量(如信任问卷)与客观指标(如评估准确率),结果表明:正确性是信任的主要驱动因素,大模型提供的解释虽提升了评估准确率,但对信任影响有限;而计划迭代优化虽未显著提高准确率,却展现出提升信任的潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for planning tasks, offering unique capabilities not found in classical planners such as generating explanations and iterative refinement. However, trust--a critical factor in the adoption of planning systems--remains underexplored in the context of LLM-based planning tasks. This study bridges this gap by comparing human trust in LLM-based planners with classical planners through a user study in a Planning Domain Definition Language (PDDL) domain. Combining subjective measures, such as trust questionnaires, with objective metrics like evaluation accuracy, our findings reveal that correctness is the primary driver of trust and performance. Explanations provided by the LLM improved evaluation accuracy but had limited impact on trust, while plan refinement showed potential for increasing trust without significantly enhancing evaluation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。