arXiv:2505.12572cs.CLcs.AI2025-05

研究超长小说生成中信息失真,找到最优压缩扩展比。

Measuring Information Distortion in Hierarchical Ultra long Novel Reconstruction:The Optimal Expansion Ratio

  • 分两阶段生成小说,分析不同压缩扩展比下的信息失真
  • 实验显示最优比例使语义失真显著降低
  • 适合关注长文本生成质量的研究者

长篇小说生成广泛采用两阶段框架(大纲→章节大纲→正文),如DOME、Plan&Write、Long Writer。然而针对超长小说(>100万字)重建中该框架的研究仍较少。基于近期文本压缩方法(LLMZip、LLM2Vec),我们开展信息论分析,量化不同压缩-扩展比率下的语义失真。通过在超长小说上的实验,发现最优压缩-扩展比率能显著减少语义失真,优于其他非最优比率。

原文摘要 · Abstract (English)

A two stage novel generation framework (outline -> section outline -> manuscript) is widely used in long novel generation,(e.g., \textsc{DOME}, \textsc{Plan\&Write}, \textsc{Long Writer}), but study of such framework in ultra long novel(>1M words) reconstruction is little. Building on recent text compression methods (\textsc{LLMZip}, \textsc{LLM2Vec}), we conduct an information-theoretic analysis to quantify semantic distortion under different compression-expansion ratios. We examine how outline length affects information preservation. Experiments on ultra-long novels show that the optimal compression-expansion ratio significantly reduces semantic distortion compared to other non-optimal compression-expansion ratio.

长文本生成信息失真压缩扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。