arXiv:2607.18257cs.HCcs.AI2026-07

用户用通用AI代理办事时,常因代理擅自行动而后悔,哪怕结果不错。

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

论文配图:Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
图 1 · 摘自论文原文
  • 按任务类型调节信任:低风险任务放手,高可见不可逆动作要确认
  • 不可逆+公开可见才引发信任暴跌,即使风险中等也如此
  • 没预览就执行,哪怕成功也会产生委托后悔感,适合设计参考

当AI代理从回答问题转向采取行动,用户面临新难题:难以预判代理的动作范围。我们称由此产生的不满为‘委托后悔’——用户后悔的不是代理出错,而是它做了自己未授权的事。在一项受控研究中,20名大学生使用通用AI代理OpenClaw完成五项日常任务,任务涵盖隐私性、风险程度与可逆性。通过5点李克特量表测量信任、感知控制、透明度、监督负担和审批偏好,并对自由文本反馈进行主题编码。结果显示:第一,参与者根据任务类型而非代理整体来调整信任,对咨询类和低风险任务给予广泛自主权,但对不可逆且外显的动作要求确认;第二,不可逆性与外部可见性共同影响信任,而非风险本身——中等风险的邮件任务导致信任最低(M = 3.10)和审批需求最高(M = 4.65),而高风险但可验证的任务未引发类似反应;第三,即使输出被评价为成功,只要代理未提供预览即执行,就会持续引发委托后悔。论文讨论了揭示动作边界、支持按任务设置自主策略、分离建议输出与执行行为的设计启示。

原文摘要 · Abstract (English)

When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. We call the resulting dissatisfaction delegation regret, a pattern in which users regret not that the agent erred, but that it acted beyond what they would have authorized. In a controlled study, 20 university students completed five common daily tasks using OpenClaw, a general-purpose AI agent, across tasks chosen to vary in privacy, stakes, and reversibility. For each task we measured trust, perceived control, transparency, supervision burden, and approval preference on 5-point Likert scales, and collected free-text reflections analyzed through thematic coding. Three findings emerged. First, participants calibrated trust per task rather than per agent: they granted wide autonomy for advisory and low-stakes tasks but demanded confirmation for irreversible, externally visible actions. Second, irreversibility combined with external visibility, rather than stakes alone, appeared to drive trust withdrawal: the moderate-stakes email task triggered the sharpest drop in trust (M = 3.10) and the highest demand for approval (M = 4.65), whereas a high-stakes but verifiable task did not produce the same response. Third, delegation regret appeared consistently when the agent executed actions without preview, even when the output was rated as successful. We discuss implications for agent designs that expose action boundaries, support per-task autonomy policies, and separate advisory output from agentic execution.

AI代理用户信任委托后悔

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。