arXiv:2509.19163cs.CL2025-09被引 7

给AI生成文本的低质量现象建立可衡量的评估框架

Measuring AI "Slop" in Text

  • 构建包含多个维度的AI文本质量分类体系
  • 发现二元判断存在主观性但与连贯性等维度相关
  • 适用于检测和偏好任务,助力理解质量判断机制

AI "slop" 是描述低质量AI生成文本的流行术语,但目前尚无统一定义或测量方法。本文通过与NLP、写作及哲学领域专家访谈,构建了"slop"的分类体系,并提出一组可解释的评估维度。通过片段级标注发现,二元是否为slop的判断具有一定主观性,但仍与连贯性、相关性等潜在维度显著相关。该框架可用于检测和二元偏好任务,有望揭示影响文本质量判断的语言与风格因素。

原文摘要 · Abstract (English)

AI "slop" is an increasingly popular term used to describe low-quality AI-generated text, but there is currently no agreed upon definition of this term nor a means to measure its occurrence. In this work, we develop a taxonomy of "slop" through interviews with experts in NLP, writing, and philosophy, and propose a set of interpretable dimensions for its assessment in text. Through span-level annotation, we find that binary "slop" judgments are (somewhat) subjective, but such determinations nonetheless correlate with latent dimensions such as coherence and relevance. Our framework can be used to evaluate AI-generated text in both detection and binary preference tasks, potentially offering new insights into the linguistic and stylistic factors that contribute to quality judgments.

文本质量AI评估自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。