arXiv:2502.01390cs.HCcs.CL2025-02中稿 · CHI 2025被引 86

研究大模型助手在日常任务中先规划后执行的效果与用户信任关系。

Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant

  • 采用先规划后执行模式,让用户全程参与决策
  • 高质计划+用户参与执行时表现最佳,但看似合理的错误计划易引发信任危机
  • 适用于设计人机协作的日常助手系统

自ChatGPT爆发以来,大型语言模型(LLMs)持续影响我们的日常生活。具备特定功能外部工具(如订票、闹钟)的大模型代理,正不断增强协助人类处理日常事务的能力。尽管大模型代理作为日常助手展现出巨大潜力,但对其基于规划与序列决策能力提供协助的理解仍有限。受近期‘人机协同规划’研究启发,本研究开展了一项实证研究(N=248),考察大模型代理在六类常见任务中的表现,涵盖不同风险等级(如机票预订、信用卡支付)。为保障用户自主性,采用‘计划-执行’模式,在模拟环境中进行分步规划与执行。分析用户在各阶段参与度对信任及协作绩效的影响。结果表明:大模型代理是把双刃剑——(1)高质量计划配合用户参与执行时表现良好;(2)看似合理的错误计划易导致用户信任下降。研究提炼出关键洞见,以校准用户信任并提升任务整体效果,对未来日常助手设计与人机协作具有重要启示。

原文摘要 · Abstract (English)

Since the explosion in popularity of ChatGPT, large language models (LLMs) have continued to impact our everyday lives. Equipped with external tools that are designed for a specific purpose (e.g., for flight booking or an alarm clock), LLM agents exercise an increasing capability to assist humans in their daily work. Although LLM agents have shown a promising blueprint as daily assistants, there is a limited understanding of how they can provide daily assistance based on planning and sequential decision making capabilities. We draw inspiration from recent work that has highlighted the value of 'LLM-modulo' setups in conjunction with humans-in-the-loop for planning tasks. We conducted an empirical study (N = 248) of LLM agents as daily assistants in six commonly occurring tasks with different levels of risk typically associated with them (e.g., flight ticket booking and credit card payments). To ensure user agency and control over the LLM agent, we adopted LLM agents in a plan-then-execute manner, wherein the agents conducted step-wise planning and step-by-step execution in a simulation environment. We analyzed how user involvement at each stage affects their trust and collaborative team performance. Our findings demonstrate that LLM agents can be a double-edged sword -- (1) they can work well when a high-quality plan and necessary user involvement in execution are available, and (2) users can easily mistrust the LLM agents with plans that seem plausible. We synthesized key insights for using LLM agents as daily assistants to calibrate user trust and achieve better overall task outcomes. Our work has important implications for the future design of daily assistants and human-AI collaboration with LLM agents.

人机协作大模型代理用户信任

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。