arXiv:2505.18128cs.CL2025-05中稿 · ACL被引 7

用随机文本片段拼接长篇故事,让AI当剪辑师而非创作者。

Frankentext: Stitching random text fragments into long-form narratives

  • AI从海量人类文本中挑选并拼接片段,仅少量修改生成连贯叙事。
  • 90%以上内容来自原文,生成故事在质量、多样性和原创性上超越普通AI写作。
  • 生成内容难辨真假,72%被顶尖检测器误判为人工撰写,引发版权争议。

我们提出Frankentexts,一种将大语言模型视为文本编排者而非作者的长篇叙事生成范式。给定写作提示和数千个随机采样的人类写作文本片段,模型需在极强约束下生成故事——例如90%以上的词元必须直接复制自提供的段落。该任务对人类而言几乎无法完成:选择与排序片段构成组合爆炸的搜索空间,而大模型隐式探索后仅做最小编辑,将选中片段缝合成连贯长篇叙事。通过大量自动与人工评估发现,Frankentexts在写作质量、多样性与原创性上显著优于常规大模型生成结果,同时保持与提示的相关性与连贯性。此外,这些生成内容对现有AI文本检测器构成严峻挑战:采用最佳Gemini 2.5 Pro配置生成的Frankentexts中,72%被当前最先进的Pangram检测器误判为人类撰写。人类评价者称赞其构思新颖、描写生动、幽默自然;但也指出长篇作品中存在突兀语气转变与段落间语法不均等问题。高质量Frankentexts的出现,引发了关于创作权与版权归属的根本性疑问:当原始素材由人类提供,而大模型负责组织成新叙事时,最终成果究竟属于谁?

原文摘要 · Abstract (English)

We introduce Frankentexts, a long-form narrative generation paradigm that treats an LLM as a composer of existing texts rather than as an author. Given a writing prompt and thousands of randomly sampled human-written snippets, the model is asked to produce a narrative under the extreme constraint that most tokens (e.g., 90%) must be copied verbatim from the provided paragraphs. This task is effectively intractable for humans: selecting and ordering snippets yields a combinatorial search space that an LLM implicitly explores, before minimally editing and stitching together selected fragments into a coherent long-form story. Despite the extreme challenge of the task, we observe through extensive automatic and human evaluation that Frankentexts significantly improve over vanilla LLM generations in terms of writing quality, diversity, and originality while remaining coherent and relevant to the prompt. Furthermore, Frankentexts pose a fundamental challenge to detectors of AI-generated text: 72% of Frankentexts produced by our best Gemini 2.5 Pro configuration are misclassified as human-written by Pangram, a state-of-the-art detector. Human annotators praise Frankentexts for their inventive premises, vivid descriptions, and dry humor; on the other hand, they identify issues with abrupt tonal shifts and uneven grammar across segments, particularly in longer pieces. The emergence of high-quality Frankentexts raises serious questions about authorship and copyright: when humans provide the raw materials and LLMs orchestrate them into new narratives, who truly owns the result?

文本生成大模型版权问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。