44%的众包文本含大模型生成内容,学界正探索应对策略。
Can Crowdsourcing Survive the LLM Era? A Community Survey on Human Data Collection

- 通过调查155位研究者,分析大模型对众包文本数据的影响
- 44%受访者发现众包数据中存在模型生成内容,93%已预料到此问题
- 当前检测手段依赖文字风格和完成速度,但应对仍不充分
大型语言模型(LLMs)广泛用作写作工具,威胁众包数据的有效性,因众包工作者可能将任务外包给模型。为更好地理解这一问题的应对方式,我们对155名自然语言处理及相关领域的研究者进行了调查,探讨其在众包自由文本收集中的经验与看法。结果显示,44%的受访者观察到其众包数据中存在大模型使用痕迹。尽管93%的人此前已预见到该问题,但一半人不清楚应采取何种预防措施。最常用的检测方法是识别独特的文本风格模式和异常快速的完成时间。总体而言,研究社区已意识到该问题并采取行动,但现有措施仍不足以完全应对。最后,本文总结出一系列指导未来在大模型时代进行众包自由文本数据采集的建议。
原文摘要 · Abstract (English)
The widespread use of Large Language Models (LLMs) as writing tools challenges the validity of crowdsourced data, as crowdworkers may outsource tasks to models. To better understand how this is addressed, we surveyed 155 researchers in NLP and related disciplines about their experiences and opinions on collecting free-text responses via crowdsourcing. This paper provides an overview of practitioners' challenges, mitigation strategies, and the foreseen implications on data quality. 44% of respondents reported observing LLM usage in their crowdsourced data. While 93% of them had anticipated this, half were unsure what precautions to take. The most prevalent detection strategies are distinctive textual style patterns and unusually fast completion times. Overall, survey responses show that the research community is aware of the problem and taking measures, but existing efforts remain insufficient to fully address it. Finally, we derive a set of considerations to guide future crowdsourced free-text data collection in the era of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。