通过叙事结构差异,93%准确区分人类与AI小说。
StoryScope: Investigating idiosyncrasies in AI fiction
- 构建10维叙事特征体系,自动分析故事深层结构。
- 仅用叙事特征达93.2%准确率,超越含风格特征模型。
- 揭示各AI模型独特叙事指纹,适合内容审核与溯源。
随着AI生成小说日益普遍,作者身份与原创性问题愈发重要。现有研究多依赖表层写作风格识别,而本文关注不依赖风格信号的叙事层面差异,如角色自主性与时间断裂。提出StoryScope管道,自动提取跨10个维度的细粒度、可解释叙事特征。应用于包含10,272个写作提示的平行语料库,每题由人类与5个大模型创作,共生成61,608篇约5,000词的故事,每篇提取304个特征。仅使用叙事特征即实现93.2% macro-F1的人类与AI检测准确率,六类作者归属达68.4% macro-F1,性能保留超97%于融合风格特征的模型。30个核心叙事特征即可捕获主要信号:AI故事过度解释主题,偏好单一主线;人类故事更强调主角道德模糊性与时间复杂性。各模型具独特叙事指纹:Claude事件推进平缓,GPT频繁使用梦境,Gemini偏重外部描写。结果显示AI故事聚集在叙事空间同一区域,而人类作品更具多样性。表明叙事构造差异可有效区分人类原创与AI生成文本。
原文摘要 · Abstract (English)
As AI-generated fiction becomes increasingly prevalent, questions of authorship and originality are becoming central to how written work is evaluated. While most existing work in this space focuses on identifying surface-level signatures of AI writing, we ask instead whether AI-generated stories can be distinguished from human ones without relying on stylistic signals, focusing on discourse-level narrative choices such as character agency and chronological discontinuity. We propose StoryScope, a pipeline that automatically induces a fine-grained, interpretable feature space of discourse-level narrative features across 10 dimensions. We apply StoryScope to a parallel corpus of 10,272 writing prompts, each written by a human author and five LLMs, yielding 61,608 stories, each ~5,000 words, and 304 extracted features per story. Narrative features alone achieve 93.2% macro-F1 for human vs. AI detection and 68.4% macro-F1 for six-way authorship attribution, retaining over 97% of the performance of models that include stylistic cues. A compact set of 30 core narrative features captures much of this signal: AI stories over-explain themes and favor tidy, single-track plots while human stories frame protagonist' choices as more morally ambiguous and have increased temporal complexity. Per-model fingerprint features enable six-way attribution: for example, Claude produces notably flat event escalation, GPT over-indexes on dream sequences, and Gemini defaults to external character description. We find that AI-generated stories cluster in a shared region of narrative space, while human-authored stories exhibit greater diversity. More broadly, these results suggest that differences in underlying narrative construction, not just writing style, can be used to separate human-written original works from AI-generated fiction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。