实验证明生成式AI在公部门能提升文档任务效率,但数据任务效果不佳。
Assessing Generative AI value in a public sector context: evidence from a field experiment
- 在真实公部门场景中开展预注册实验,对比使用与未使用Gen AI的人员表现。
- 处理文档任务时,答案质量提升17%,完成时间缩短34%;数据任务质量下降12%。
- 结果表明Gen AI效果因任务类型而异,适合评估其在实际工作中的适用性。
生成式AI(Gen AI)的兴起引发了对其能否提升各类任务生产力的关注。本文在公共部门背景下,通过一项预注册实验,研究了Gen AI对复杂知识型任务的影响。在建立基线后,针对两类复合任务——文档理解与数据分析——发现混合结果:在文档任务中,使用Gen AI的实验组答案质量评分提升17%(人工评估),任务完成时间缩短34%;而在数据任务中,实验组质量评分下降12%,完成时间无显著差异。结果表明,Gen AI的效益可能依赖于具体任务类型及执行者。研究还结合现场观察、事后问卷及反馈研讨会,提供了实践启示。
原文摘要 · Abstract (English)
The emergence of Generative AI (Gen AI) has motivated an interest in understanding how it could be used to enhance productivity across various tasks. We add to research results for the performance impact of Gen AI on complex knowledge-based tasks in a public sector setting. In a pre-registered experiment, after establishing a baseline level of performance, we find mixed evidence for two types of composite tasks related to document understanding and data analysis. For the Documents task, the treatment group using Gen AI had a 17% improvement in answer quality scores (as judged by human evaluators) and a 34% improvement in task completion time compared to a control group. For the Data task, we find the Gen AI treatment group experienced a 12% reduction in quality scores and no significant difference in mean completion time compared to the control group. These results suggest that the benefits of Gen AI may be task and potentially respondent dependent. We also discuss field notes and lessons learned, as well as supplementary insights from a post-trial survey and feedback workshop with participants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。