研究发现人类在使用大模型时,反而被其建议影响,效率提升但分析可能偏颇。
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
- 通过人机协作实验,对比人类与大模型生成主题的差异。
- 大模型可快速生成高重合度主题,但遗漏部分文档特异性内容。
- 适合关注人机协作中偏见风险的研究者与实践者。
大型语言模型(LLMs)在多种分析任务中表现出接近人类的性能,因此被用于节省时间和人力密集型分析。然而,它们在政策研究等高度专业化和开放性任务中的能力仍存疑。本文通过结构化用户研究,考察了人-大模型协作在专题发现与专题分配两个阶段的效率与准确性。研究将大模型建议与专家标注结合,观察其对传统仅由人类完成分析的影响。结果显示,大模型生成的主题列表与人类生成列表有显著重合,但在捕捉文档特定主题方面存在细微缺陷。尽管大模型建议能显著提升任务完成速度,但也可能引入锚定偏差,影响分析的深度与细致程度,引发关于效率提升与偏见风险之间权衡的深刻质疑。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown capabilities close to human performance in various analytical tasks, leading researchers to use them for time and labor-intensive analyses. However, their capability to handle highly specialized and open-ended tasks in domains like policy studies remains in question. This paper investigates the efficiency and accuracy of LLMs in specialized tasks through a structured user study focusing on Human-LLM partnership. The study, conducted in two stages-Topic Discovery and Topic Assignment-integrates LLMs with expert annotators to observe the impact of LLM suggestions on what is usually human-only analysis. Results indicate that LLM-generated topic lists have significant overlap with human generated topic lists, with minor hiccups in missing document-specific topics. However, LLM suggestions may significantly improve task completion speed, but at the same time introduce anchoring bias, potentially affecting the depth and nuance of the analysis, raising a critical question about the trade-off between increased efficiency and the risk of biased analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。