arXiv:2603.27768cs.CL2026-03

首个针对长尾实体文本生成的多语言评测基准,揭示模型对罕见实体的表达偏差。

TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities

  • 构建覆盖三语的长尾实体评测集TailNLG,基于Wikidata数据
  • 模型在罕见实体上的生成得分更低,不确定性更高,存在系统性偏差
  • 适用于知识图谱、多语言生成研究者,推动更公平的评估框架

结构化知识的自动文本化是使知识图谱对非专家用户可访问并支持检索增强生成的关键任务。尽管近期数据到文本生成在多语言覆盖上取得进展,但对罕见实体(即长尾实体)表达偏见的关注仍不足。本文首次系统研究数据到文本生成中的长尾实体问题,提出新多语言基准TailNLG,涵盖英语、意大利语和西班牙语,基于Wikidata构建,覆盖不同流行度的实体。我们在零样本设置下评估三类大语言模型在稀有与常见实体上的表现,并与现有WebNLG基准对比。结果表明,模型对长尾实体存在一致偏差:嵌入得分更低,不确定性更高。此外,长尾影响因模型和语言而异,现有评估指标未能一致捕捉差异,凸显了更可靠评估框架的必要性。

原文摘要 · Abstract (English)

The automatic verbalization of structured knowledge is a key task for making knowledge graphs accessible to non-expert users and supporting retrieval-augmented generation systems. Although recent advances in Data-to-Text generation have improved multilingual coverage, little attention has been paid to potential biases in the verbalization of rare entities, frequently known as long-tail entities. In this work, we present the first systematic study of long-tail entities in Data-to-Text generation. We introduce TailNLG, a new multilingual benchmark in English, Italian, and Spanish, built from Wikidata and covering entities with varying levels of popularity. We evaluate three different families of large language models in zero-shot settings and compare their performance on rare versus common entities, as well as against the established WebNLG benchmark. Our results reveal a consistent bias against long-tail entities: embedding-based scores are lower, and model uncertainty is higher for rare entities. We further show that the impact of long-tail entities varies across models and languages, and that existing evaluation metrics do not consistently capture these differences, highlighting the need for more reliable evaluation frameworks.

知识图谱多语言长尾实体文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。