LLMs在群体观点预测中可优于人类,因误差更小、偏差更独立。
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
- 将观点理解视为群体判断的潜在估计,构建评估框架。
- 实测显示LLMs在主观任务中群体意见预测误差更低。
- 适合需要大规模群体意见建模的研究与应用。
尽管大型语言模型(LLMs)被广泛用于大规模标注,但通常被视为权宜之计而非人类观点的忠实估计。本文挑战这一假设:将观点采纳视为潜在群体层面判断的估计,分析现代LLMs在预测主观任务中子群体集体意见时超越人类标注者(包括同群人类)的条件,并证明这些条件在实践中普遍存在。该优势源于LLMs作为估计器的结构性特征,包括低方差和表示与处理偏差间的弱耦合,而非任何生活经验。研究识别出明确的场景,使LLMs成为统计上更优的一线估计工具,同时也明确了人类判断仍不可或缺的边界。这一发现将LLMs从成本节约的妥协方案,重新定位为系统性估算集体人类观点的可靠工具。
原文摘要 · Abstract (English)
Although large language models (LLMs) are increasingly used as annotators at scale, they are typically treated as a pragmatic fallback rather than a faithful estimator of human perspectives. This work challenges that presumption. By framing perspective-taking as the estimation of a latent group-level judgment, we characterize the conditions under which modern LLMs can outperform human annotators, including in-group humans, when predicting aggregate subgroup opinions on subjective tasks, and show that these conditions are common in practice. This advantage arises from structural properties of LLMs as estimators, including low variance and reduced coupling between representation and processing biases, rather than any claim of lived experience. Our analysis identifies clear regimes where LLMs act as statistically superior frontline estimators, as well as principled limits where human judgment remains essential. These findings reposition LLMs from a cost-saving compromise to a principled tool for estimating collective human perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。