arXiv:2510.04400cs.CLcs.AI2025-10被引 1

大模型续写故事时能保持语义关系一致性,值得一看。

Large Language Models Preserve Semantic Isotopies in Story Continuations

  • 用1万条故事开头测试5个大模型续写能力
  • 在限定词数内,语义关系结构基本不变
  • 适合研究语言模型语义保真度的学者

本文探究大型语言模型(LLMs)生成文本中的语义特性,扩展了分布语义与结构语义之间关联的先前认知。我们通过使用10,000个ROCStories故事开头,由五个不同LLM完成续写,检验其是否保持语义同构性。首先验证GPT-4o从语言基准中提取同构关系的能力,随后将其应用于生成文本。我们分析了同构性的结构属性(覆盖率、密度、扩散范围)与语义特征,评估续写对它们的影响。结果显示,在给定的词元范围内,LLM的续写能够保持语义同构性于多个维度。

原文摘要 · Abstract (English)

In this work, we explore the relevance of textual semantics to Large Language Models (LLMs), extending previous insights into the connection between distributional semantics and structural semantics. We investigate whether LLM-generated texts preserve semantic isotopies. We design a story continuation experiment using 10,000 ROCStories prompts completed by five LLMs. We first validate GPT-4o's ability to extract isotopies from a linguistic benchmark, then apply it to the generated stories. We then analyze structural (coverage, density, spread) and semantic properties of isotopies to assess how they are affected by completion. Results show that LLM completion within a given token horizon preserves semantic isotopies across multiple properties.

大模型语义保持故事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。