提出新指标衡量用户对AI的依赖程度,基于反事实工作流计算认知负担转移量。
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

- 通过构建无AI时的反事实任务流程,计算用户节省的任务步骤比例。
- 在限时条件下,依赖度提升43%(p=0.018),传统指标无法捕捉此变化。
- 可帮助开发者判断何时依赖合理,适合评估AI协作中的责任边界。
AI工具日益融入真实工作流程。现有依赖度测量方法多聚焦于输出采纳或自我报告,未反映用户与工具间任务努力的分配情况。本文提出‘离负荷得分’(Offloading Score),一种基于模拟的依赖度度量,通过估算无工具时用户完成任务所需步骤,计算使用工具后节省的步骤比例。在40名开发者的受控实验中,我们验证了该指标的有效性,并考察时间压力下依赖度的变化。结果表明,在时间约束下,离负荷得分显著上升43%(p=0.018),而基于使用频率和自评的基准度量则无差异。进一步发现,高依赖表现为更多子任务委托及更直接复用AI输出。最后,我们展示将离负荷得分与任务目标(如代码理解)结合,可识别依赖是否恰当。本框架提供两个价值:一是用户自我评估依赖程度的工具,二是设计师用于缓解过度依赖的量化信号。
原文摘要 · Abstract (English)
AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between users and tools. Here, we introduce offloading score, a measure of reliance that quantifies the fraction of cognitive effort offloaded to an AI tool. Offloading Score is simulation-based -- we construct a counterfactual workflow by estimating how the user would have completed the task without the tool, and then computing the fraction of steps saved by using the tool. We validate offloading score through intrinsic evaluations of metric validity, and a controlled user study ($n=40$) with developers performing programming tasks using AI tools. We vary time pressure to test whether reliance measures capture the known increase in reliance under time pressure. We show that offloading score detects significantly higher reliance in time-constrained settings ($+43\%$, $p=0.018$), while usage-based and self-reported baseline measures of reliance do not distinguish the conditions. We complement this with descriptive insights showing that higher reliance manifests as greater delegation of subtasks to the tool and more direct reuse of AI outputs. Finally, we demonstrate an approach of using offloading score in combination with target outcomes of a task (e.g., code understanding) to identify when reliance may be (in)appropriate. Our framework offers two contributions: an instrument users can apply to measure and reflect on their own reliance, and a quantitative signal that agent designers can utilize to mitigate overreliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。