用新闻检索生成芬兰语讽刺定义,提升政治相关性但幽默感有限
Grounded Satirical Generation with RAG
- 基于新闻检索增强生成,结合芬兰语语境生成讽刺词义
- 生成内容更显政治性而非幽默性,检索提升政治相关性
- 提供可复现数据集与评测框架,适合跨文化幽默研究
幽默生成对大语言模型仍是挑战,因其主观性强。本文聚焦讽刺这种高度依赖语境的幽默形式,提出一种基于当前新闻的检索增强生成(RAG)新管道,用于生成芬兰语背景下的讽刺词典释义。我们构建了面向该任务的评估框架,并由六名人工标注者对100条生成结果进行标注,支持在文化背景、词源类型及RAG使用与否等多重条件下分析。结果显示,生成内容被感知为更具政治性而非幽默性;基于主题的词选择和RAG均提升了输出的政治相关性,但未显著提升幽默效果。此外,五种主流模型作为评判者,在政治相关性上与人类判断有良好相关性,但在幽默性判断上表现不佳。代码与标注数据集已开源,以支持后续研究。
原文摘要 · Abstract (English)
Humor generation remains challenging task for Large Language Models (LLMs), due to their subjective nature. We focus on satire, a form of humor strongly shaped by context. In this work, we present a novel pipeline for grounded satire generation that uses Retrieval-Augmented Generation (RAG) over current news to produce satirical dictionary definitions in the Finnish context. We also introduce a new task-specific evaluation framework and annotate 100 generated definitions with six human annotators, enabling analysis across multiple experimental conditions, including cultural background, source-word type, and the presence or absence of RAG. Our results show that the generated definitions are perceived as more political than humorous. Both topic-based word selection and RAG improve the political relevance of the outputs, but neither yields clear gains in humor generation. In addition, our LLM-as-a-judge evaluation of five state-of-the-art models indicates that LLMs correlate well with human judgments on political relevance, but perform poorly on humor. We release our code and annotated dataset to support further research on grounded satire generation and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。