AI写作评价差异源于读者偏好不同,非文本质量本身高低。
The Reader is the Metric: How Textual Features and Reader Profiles Explain Conflicting Evaluations of AI Creative Writing
- 用17个无参考文本特征分析文学性,建模读者偏好向量。
- 发现两类读者:重可读性的非专家与重主题的专家,评价标准迥异。
- 提出应根据读者类型设计评估框架,适合研究AI创作评价的学者。
近期比较AI生成与人类创作文学文本的研究得出矛盾结论:部分认为AI已超越人类,另一些则认为仍显不足。我们假设这些分歧主要源于读者对文学的理解与价值判断差异,而非文本内在质量。基于五个公开数据集(1,471篇故事,101名标注者,包括评论家、学生和普通读者),我们(i)提取17个无参考文本特征(如连贯性、情感波动、平均句长等);(ii)建模个体读者偏好,得到反映其文本优先级的特征重要性向量;(iii)在共享“偏好空间”中分析这些向量。结果发现读者向量聚为两类:'表面聚焦型'(多为非专业人士),重视可读性与文本丰富性;'整体聚焦型'(多为专业人士),关注主题发展、修辞多样性与情感动态。研究定量揭示了文学质量评估依赖于文本特征与读者偏好的匹配度,呼吁建立读者敏感型评估框架。
原文摘要 · Abstract (English)
Recent studies comparing AI-generated and human-authored literary texts have produced conflicting results: some suggest AI already surpasses human quality, while others argue it still falls short. We start from the hypothesis that such divergences can be largely explained by genuine differences in how readers interpret and value literature, rather than by an intrinsic quality of the texts evaluated. Using five public datasets (1,471 stories, 101 annotators including critics, students, and lay readers), we (i) extract 17 reference-less textual features (e.g., coherence, emotional variance, average sentence length...); (ii) model individual reader preferences, deriving feature importance vectors that reflect their textual priorities; and (iii) analyze these vectors in a shared "preference space". Reader vectors cluster into two profiles: 'surface-focused readers' (mainly non-experts), who prioritize readability and textual richness; and 'holistic readers' (mainly experts), who value thematic development, rhetorical variety, and sentiment dynamics. Our results quantitatively explain how measurements of literary quality are a function of how text features align with each reader's preferences. These findings advocate for reader-sensitive evaluation frameworks in the field of creative text generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。