arXiv:2502.07732cs.CYcs.AI2025-02ICML被引 4

AI时代人类数据质量下滑,需用内在动机替代外部激励。

When Incentives Backfire, Data Stops Being Human

  • 用内在动机替代外部奖励,重构数据采集系统。
  • 当前系统重速度效率,致参与度与数据质量双降。
  • 适合关注可持续数据生态的研究者与平台设计者。

人工智能的进步依赖于人类生成的数据,包括标注者市场和互联网上的广泛内容。然而,大语言模型的广泛应用正威胁这些平台上人类生成数据的质量与完整性。我们认为,这一问题不仅在于过滤生成内容,更暴露了数据收集系统设计中的深层缺陷:现有系统往往以速度、规模和效率为优先,牺牲了人类的内在动机,导致参与度下降和数据质量恶化。我们提出,重新设计数据收集系统,使其与贡献者的内在动机对齐,而非仅依赖外部激励,可在保持贡献者信任和长期参与的前提下,实现高质量数据的规模化获取。

原文摘要 · Abstract (English)

Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality and integrity of human-generated data on these very platforms. We argue that this issue goes beyond the immediate challenge of filtering AI-generated content -- it reveals deeper flaws in how data collection systems are designed. Existing systems often prioritize speed, scale, and efficiency at the cost of intrinsic human motivation, leading to declining engagement and data quality. We propose that rethinking data collection systems to align with contributors' intrinsic motivations -- rather than relying solely on external incentives -- can help sustain high-quality data sourcing at scale while maintaining contributor trust and long-term participation.

数据质量激励机制人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。