arXiv:2510.13939cs.CLcs.AI2025-10被引 12

用作者作品微调AI,生成文本比专业写手更受读者欢迎。

Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers

  • 用作者完整作品微调AI模型,让其模仿特定作家风格。
  • 微调后AI文本被读者认为更贴近原作者文风且质量更高。
  • 微调成本仅81美元/作者,远低于人类作者报酬,适合版权争议场景。

使用受版权保护的书籍训练AI引发了作者诉讼,但这类模型能否生成高质量、具作者风格的文学文本尚不明确。我们开展了一项预注册研究,比较了28名MFA背景写作者与3个前沿模型(ChatGPT、Claude、Gemini)在450字以内的文本,模仿50位获奖作家的风格。盲测中,MFA读者对仅通过上下文提示生成的AI文本在风格契合度(OR=0.16)和质量(OR=0.13)上均显著偏好人类写作;而普通读者无风格偏好(OR=1.06),但更倾向AI生成内容(质量OR=1.82)。当对ChatGPT进行作者作品微调后,结果反转:MFA读者更青睐AI在风格(OR=8.16)和质量(OR=1.87)上的表现,普通读者偏好更强(风格OR=16.65;质量OR=5.42)。两种读者群体均更偏爱微调后的AI,但读者类型与偏好存在显著交互作用(风格p=0.021,质量p<10^-4)。微调输出极少被检测为AI生成(3%),而提示生成的达97%。中介分析表明,微调消除了可被检测的AI特征,改变了可检测性与偏好之间的关系。尽管未考虑将AI输出转化为出版级文本所需投入,但平均微调成本仅81美元/作者,较典型写作者报酬降低99.7%。作者专属微调实现了非抄录式文本生成,且广受读者欢迎,为版权法第四条合理使用提供了实证支持。

原文摘要 · Abstract (English)

The use of copyrighted books for training AI has sparked lawsuits from authors concerned about AI generating derivative content. Yet whether these models can produce high-quality literary text emulating authors' voices remains unclear. We conducted a preregistered study comparing MFA-trained writers with three frontier models (ChatGPT, Claude, Gemini) writing up to 450-word excerpts emulating 50 award-winning authors' styles. In blind pairwise evaluations by 28 MFA-trained readers and 516 college-educated general readers, AI text from in-context prompting was strongly disfavored by MFA readers for stylistic fidelity (OR=0.16) and quality (OR=0.13), while general readers showed no fidelity preference (OR=1.06) but favored AI for quality (OR=1.82). Fine-tuning ChatGPT on authors' complete works reversed these results: MFA readers favored AI for fidelity (OR=8.16) and quality (OR=1.87), with general readers showing even stronger preference (fidelity OR=16.65; quality OR=5.42). Both groups preferred fine-tuned AI, but the writer-type X reader-type interaction remained significant (p=0.021 for fidelity; p<10^-4 for quality), indicating general readers favored AI by a wider margin. Effects are robust under cluster-robust inference and generalize across authors in heterogeneity analyses. Fine-tuned outputs were rarely flagged as AI-generated (3% vs. 97% for prompting) by leading detectors. Mediation analysis shows fine-tuning eliminates detectable AI quirks that penalize in-context outputs, altering the nexus between detectability and preference. While not accounting for effort to transform AI output into publishable prose, the median fine-tuning cost of $81 per author represents a 99.7% reduction versus typical writer compensation. Author-specific fine-tuning enables non-verbatim AI writing preferred over expert human writing, providing evidence relevant to copyright's fourth fair-use factor.

AI写作版权争议微调读者偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。