让AI在表格中每一步都可审计可控,用户能实时干预决策过程。
Auditing and Controlling AI Agent Actions in Spreadsheets

- 将AI执行拆解为可审计的步骤,每步允许用户查看和干预。
- 16人实验显示,实时参与使用户更准确发现错误并理解任务。
- 适合需要高透明度与人机协作的办公自动化场景。
AI代理在知识工作中的能力快速提升,但用户难以有效监督其执行过程。代理可自主完成复杂多步任务,但中间推理和输出大量堆积,用户只能在最后接收结果,无法介入决策。这种不可见性导致用户无法检查假设、发现错误或纠正偏差。尤其在表格环境中,每个决策直接记录在用户数据中,影响深远。我们提出Pista,一个将执行过程分解为可审计、可控制动作的表格AI代理。初步研究(N=8)和对比实验(N=16)表明,主动参与执行不仅改善任务结果,还提升用户对任务的理解、对代理的信任感以及在流程中的角色认同。用户能识别自身意图是否被体现,发现事后审查难以察觉的错误,并产生对产出的共担感。研究证明,真正有效的监督不是改进事后审查,而是实现在决策过程中的人类参与。
原文摘要 · Abstract (English)
Advances in AI agent capabilities have outpaced users' ability to meaningfully oversee their execution. AI agents can perform sophisticated, multi-step knowledge work autonomously from start to finish, yet this process remains effectively inaccessible during execution, often buried within large volumes of intermediate reasoning and outputs: by the time users receive the output, all underlying decisions have already been made without their involvement. This lack of transparency leaves users unable to examine the agent's assumptions, identify errors before they propagate, or redirect execution when it deviates from their intent. The stakes are particularly high in spreadsheet environments, where process and artifact are inseparable. Each decision the agent makes is recorded directly in cells that belong to and reflect on the user. We introduce Pista, a spreadsheet AI agent that decomposes execution into auditable, controllable actions, providing users with visibility into the agent's decision-making process and the capacity to intervene at each step. A formative study (N = 8) and a within-subjects summative evaluation (N = 16) comparing Pista to a baseline agent demonstrated that active participation in execution influenced not only task outcomes but also users' comprehension of the task, their perception of the agent, and their sense of role within the workflow. Users identified their own intent reflected in the agent's actions, detected errors that post-hoc review would have failed to surface, and reported a sense of co-ownership over the resulting output. These findings indicate that meaningful human oversight of AI agents in knowledge work requires not improved post-hoc review mechanisms, but active participation in decisions as they are made.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。