无需真实答案,用答题相关性检测AI作弊的标注数据质量评估方法
Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth
- 基于工作者答题相关性,不依赖真实标签进行评分
- 在真实数据集上有效识别低努力水平的AI代答行为
- 适用于多选标注等无标准答案的任务场景
生成式AI的兴起凸显了高质量人工反馈在构建可信AI系统中的关键作用。然而,众包工人越来越多地使用大语言模型(LLMs)生成回答,导致本应反映人类意见的数据集可能被污染。现有LLM检测方法通常依赖高维训练数据(如文本),难以适用于多选标注等标注任务。本文研究了同行预测(peer prediction)机制在无真实标签条件下检测众包中由LLM辅助作弊的潜力,重点关注标注任务。我们的方法在已知请求者提供的部分LLM生成标签条件下,量化工作者回答间的相关性。基于前期研究,提出一种无需训练的评分机制,并在考虑LLM串通的众包模型下提供理论保证。我们建立了该方法有效的条件,并在真实世界众包数据集上实证验证其对低努力作弊行为的鲁棒检测能力。
原文摘要 · Abstract (English)
The recent success of generative AI highlights the crucial role of high-quality human feedback in building trustworthy AI systems. However, the increasing use of large language models (LLMs) by crowdsourcing workers poses a significant challenge: datasets intended to reflect human input may be compromised by LLM-generated responses. Existing LLM detection approaches often rely on high-dimensional training data such as text, making them unsuitable for annotation tasks like multiple-choice labeling. In this work, we investigate the potential of peer prediction -- a mechanism that evaluates the information within workers' responses without using ground truth -- to mitigate LLM-assisted cheating in crowdsourcing with a focus on annotation tasks. Our approach quantifies the correlations between worker answers while conditioning on (a subset of) LLM-generated labels available to the requester. Building on prior research, we propose a training-free scoring mechanism with theoretical guarantees under a crowdsourcing model that accounts for LLM collusion. We establish conditions under which our method is effective and empirically demonstrate its robustness in detecting low-effort cheating on real-world crowdsourcing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。