研究如何让人类高效验证智能体行为,提出新界面减少纠错时间。
Overseeing Agents Without Constant Oversight: Challenges and Opportunities
- 设计新界面展示智能体推理过程,降低用户理解负担。
- 实验显示纠错时间减少,但准确率未明显提升。
- 适合关注人机协作验证的AI系统设计者阅读。
为实现人类对智能体系统的监督,现有方法通常提供推理与行动步骤的记录。如何在信息量充足与不冗余之间取得平衡仍是关键挑战。通过三个用户研究,我们在计算机用户代理任务中评估了基础动作记录的实用性,探索三种替代设计方案,并测试了一种新型界面在问答任务中发现错误的效果。结果表明,当前实践操作繁琐,效率受限。而我们提出的方案显著减少了用户发现错误所需时间。尽管参与者对决策信心提升,但最终准确率并未显著改善。本研究揭示了人类验证智能体系统所面临的挑战:包括处理内置假设、用户主观且动态变化的正确性标准,以及传达智能体行为过程的局限性与重要性。
原文摘要 · Abstract (English)
To enable human oversight, agentic AI systems often provide a trace of reasoning and action steps. Designing traces to have an informative, but not overwhelming, level of detail remains a critical challenge. In three user studies on a Computer User Agent, we investigate the utility of basic action traces for verification, explore three alternatives via design probes, and test a novel interface's impact on error finding in question-answering tasks. As expected, we find that current practices are cumbersome, limiting their efficacy. Conversely, our proposed design reduced the time participants spent finding errors. However, although participants reported higher levels of confidence in their decisions, their final accuracy was not meaningfully improved. To this end, our study surfaces challenges for human verification of agentic systems, including managing built-in assumptions, users' subjective and changing correctness criteria, and the shortcomings, yet importance, of communicating the agent's process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。