arXiv:2504.06549cs.CYcs.AI2025-04被引 4

为评估AI创作内容的社会影响,需建立真实场景下的基准测试。

Societal Impacts Research Requires Benchmarks for Creative Composition Tasks

  • 基于200万用户提问分析,发现创作类任务是常见使用场景。
  • 现有基准与真实创作需求不匹配,可能掩盖潜在风险。
  • 适合关注AI社会影响的研究者和政策制定者阅读。

能够自动化认知任务的基础模型代表了重大的技术变革,但其社会影响仍不明确。这些系统虽有望带来重大进展,但也可能使信息生态系统充斥着程式化、同质化且可能具有误导性的合成内容。因此,开发基于实际应用场景的基准测试至关重要,尤其在风险最显著的领域。通过对200万条语言模型用户提示进行主题分析,我们识别出创意写作任务是用户寻求日常创造性帮助的普遍使用类别。细粒度分析揭示了当前基准与这类任务使用模式之间的不匹配。关键在于,那些目前缺乏充分评估的使用场景可能导致负面下游影响。本文主张,针对创意创作任务建立基准是理解AI生成内容社会危害的必要步骤。我们呼吁提高使用模式的透明度,以指导新基准的开发,从而有效衡量模型在创造力方面的进展及其影响。

原文摘要 · Abstract (English)

Foundation models that are capable of automating cognitive tasks represent a pivotal technological shift, yet their societal implications remain unclear. These systems promise exciting advances, yet they also risk flooding our information ecosystem with formulaic, homogeneous, and potentially misleading synthetic content. Developing benchmarks grounded in real use cases where these risks are most significant is therefore critical. Through a thematic analysis using 2 million language model user prompts, we identify creative composition tasks as a prevalent usage category where users seek help with personal tasks that require everyday creativity. Our fine-grained analysis identifies mismatches between current benchmarks and usage patterns among these tasks. Crucially, we argue that the same use cases that currently lack thorough evaluations can lead to negative downstream impacts. This position paper argues that benchmarks focused on creative composition tasks is a necessary step towards understanding the societal harms of AI-generated content. We call for greater transparency in usage patterns to inform the development of new benchmarks that can effectively measure both the progress and the impacts of models with creative capabilities.

AI社会影响基准测试创意生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。