arXiv:2505.23166cs.CL2025-05ACL被引 2

用语言模型重述文学文本,提炼深层主题

Tell, Don't Show: Leveraging Language Models' Abstractive Retellings to Model Literary Themes

  • 让小模型重述文学片段,把细节描写转为抽象主题
  • 相比传统方法,主题更精准且信息量更高
  • 适合研究文学主题与文化认同的学者和教育者

传统的词袋主题建模方法(如LDA)在处理文学文本时表现不佳。文学语言注重沉浸式感官细节,而非抽象描述或说明,即所谓‘展示而非讲述’。我们提出Retell方法,通过轻量级生成式语言模型对文本片段进行重述,将表面叙事转化为高层概念与主题。再对重述结果运行LDA,得到比直接使用LDA或让模型直接列出主题更精确、更丰富的主题。在一项关于高中英语教材中种族/文化身份的主题研究中,我们对比了该方法输出与专家标注的结果,验证了其在文化分析中的潜力。

原文摘要 · Abstract (English)

Conventional bag-of-words approaches for topic modeling, like latent Dirichlet allocation (LDA), struggle with literary text. Literature challenges lexical methods because narrative language focuses on immersive sensory details instead of abstractive description or exposition: writers are advised to "show, don't tell." We propose Retell, a simple, accessible topic modeling approach for literature. Here, we prompt resource-efficient, generative language models (LMs) to tell what passages show, thereby translating narratives' surface forms into higher-level concepts and themes. By running LDA on LMs' retellings of passages, we can obtain more precise and informative topics than by running LDA alone or by directly asking LMs to list topics. To investigate the potential of our method for cultural analytics, we compare our method's outputs to expert-guided annotations in a case study on racial/cultural identity in high school English language arts books.

主题建模文学分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。