arXiv:2602.13283cs.AIcs.CY2026-02综述

工作场景中人们对AI准确性的要求远高于个人使用。

Accuracy Standards for AI at Work vs. Personal Life: Evidence from an Online Survey

  • 区分工作与个人场景,定义上下文相关的准确性标准
  • 工作中高精度需求占比24.1%,个人仅8.8%,差距显著
  • 工具不可用时,个人生活受影响更大,适合职场研究者参考

我们研究了人们在工作与个人场景中使用AI工具时对准确性的权衡,以及这种权衡的决定因素和当AI应用不可用时的应对方式。由于现代AI系统(尤其是生成模型)常产生可接受但不完全一致的输出,我们将“准确性”定义为情境相关的可靠性:输出与用户意图的契合程度,在容差阈值内,该阈值取决于风险水平和修正成本。基于一项在线调查(N=300),在有准确度数据的受访者中(N=170),工作中要求高准确性的比例为24.1%,个人生活中仅为8.8%(高出15.3个百分点;z=6.29,p<0.001)。在更宽泛的前两档定义下(67.0% vs. 32.9%)及全量表(1-5分制,均值分别为3.86和3.08)中,差异依然明显。重度使用和经验丰富的用户在工作场景中表现出更严格的准确性标准(H2)。当工具不可用时(H3),受访者报告个人日常受到更大干扰(34.1%对比15.3%,p<0.01)。主文本聚焦核心发现,测试分类与统计功效推导置于技术附录。

原文摘要 · Abstract (English)

We study how people trade off accuracy when using AI-powered tools in professional versus personal contexts for adoption purposes, the determinants of those trade-offs, and how users cope when AI/apps are unavailable. Because modern AI systems (especially generative models) can produce acceptable but non-identical outputs, we define "accuracy" as context-specific reliability: the degree to which an output aligns with the user's intent within a tolerance threshold that depends on stakes and the cost of correction. In an online survey (N=300), among respondents with both accuracy items (N=170), the share requiring high accuracy (top-box) is 24.1% at work vs. 8.8% in personal life (+15.3 pp; z=6.29, p<0.001). The gap remains large under a broader top-two-box definition (67.0% vs. 32.9%) and on the full 1-5 ordinal scale (mean 3.86 vs. 3.08). Heavy app use and experience patterns correlate with stricter work standards (H2). When tools are unavailable (H3), respondents report more disruption in personal routines than at work (34.1% vs. 15.3%, p<0.01). We keep the main text focused on these substantive results and place test taxonomy and power derivations in a technical appendix.

AI使用用户体验准确性标准行为研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。