arXiv:2606.12599cs.CL2026-06

用波斯谚语生成故事,测试大模型的抽象到具体理解能力

Constrained Semantic Decompression in LLMs through Persian Proverb-Conditioned Story Generation

  • 以波斯谚语为约束,生成符合道德和因果逻辑的故事
  • 发现当前大模型存在显著的语义压缩鸿沟,表面流畅但内涵失真
  • 显式推理与迭代优化可部分缓解问题,适合文化知识研究者参考

将一个密集而抽象的谚语转化为引人入胜且道德忠实的叙事,需要深厚的跨文化理解与坚实的语义基础。我们将其定义为一项受约束的语义解压缩任务,并以谚语条件下的故事生成作为大语言模型(LLMs)从抽象到具体实现的测试基准。聚焦波斯语,我们构建了谚语对齐叙事数据集(PAND),包含人类撰写的配对故事与明确含义。通过结合人工校准的基于LLM的评判器与结构化指标的混合评估框架,我们分析了多种提示策略下模型的表现。研究发现,当前大模型普遍存在解压缩鸿沟:虽能实现表面流畅,却难以忠实还原谚语中的深层道德与因果结构。进一步表明,显式推理与迭代修正可在一定程度上缓解此类失败,暗示多数解压缩错误源于抽象意义向叙事形式转化的困难,而非缺乏相关知识。该任务可自然拓展至其他形式的文化知识压缩。

原文摘要 · Abstract (English)

Transforming a dense, abstract proverb into an engaging and morally faithful narrative requires deep cultural understanding and robust semantic grounding. We frame this problem as a \emph{constrained semantic decompression} task and study proverb-conditioned story generation as a testbed for abstraction-to-realization in large language models (LLMs). Focusing on Persian, we introduce the Proverb Aligned Narrative Dataset (PAND), pairing proverbs with human-written stories and explicit meanings. By a hybrid evaluation framework that combines human-calibrated LLM-as-a-Judge with structural metrics, we analyze model behavior across multiple prompting regimes. Our findings reveal a persistent \emph{decompression gap}: current LLMs often achieve strong surface-level fluency while failing to faithfully instantiate the underlying moral and causal structure encoded in proverbs. We further show that explicit reasoning and iterative refinement can partially mitigate these failures, suggesting that many decompression errors arise from difficulties in translating abstract meaning into narrative form rather than a complete lack of relevant knowledge. Our proposed task naturally extends to other forms of compressed cultural knowledge.

语义解压缩文化知识大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。