研究诗歌中空白符的分布,揭示诗人创作意图与大模型生成差异
so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
- 分析1.9万首诗的空白符使用模式,发现其反映艺术风格与形式选择
- 对比5.1万条大模型生成诗与1.2万条未发表诗,发现空白符分布显著不同
- 强调空白符处理对预训练数据构建的影响,适合关注文本格式与模型偏见的研究者
空白符是诗歌形式的关键组成部分,既体现对规范形式的遵循,也反映对其的反叛。每首诗的空白符分布反映了诗人的艺术选择,是诗歌语义与空间结构的重要特征。尽管诗歌既是悠久的艺术形式,也是大语言模型(LLMs)的生成任务,但自然语言处理领域对空白符的关注仍不足。本研究基于来自Poetry Foundation的1.9万首英文公开诗歌,分析了4000位诗人的空白符使用情况。我们发布其中2800首保留格式的公共领域诗歌,以促进该领域研究。将这些已发表诗歌与5.1万条LLM生成诗歌、1.2万条在线社区未发表诗歌进行比较,探讨空白符在不同时期、诗体及数据源中的分布差异。此外,我们发现不同的文本处理方式会导致诗歌空白符表示显著不同,由此呼吁重新审视用于构建预训练数据集的处理策略。
原文摘要 · Abstract (English)
Whitespace is a critical component of poetic form, reflecting both adherence to standardized forms and rebellion against those forms. Each poem's whitespace distribution reflects the artistic choices of the poet and is an integral semantic and spatial feature of the poem. Yet, despite the popularity of poetry as both a long-standing art form and as a generation task for large language models (LLMs), whitespace has not received sufficient attention from the NLP community. Using a corpus of 19k English-language published poems from Poetry Foundation, we investigate how 4k poets have used whitespace in their works. We release a subset of 2.8k public-domain poems with preserved formatting to facilitate further research in this area. We compare whitespace usage in the published poems to (1) 51k LLM-generated poems, and (2) 12k unpublished poems posted in an online community. We also explore whitespace usage across time periods, poetic forms, and data sources. Additionally, we find that different text processing methods can result in significantly different representations of whitespace in poetry data, motivating us to use these poems and whitespace patterns to discuss implications for the processing strategies used to assemble pretraining datasets for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。