arXiv:2604.00019cs.CLcs.AI2026-04中稿 · LREC 2026

构建可控流行度的多语言长文本事实性评测数据集

The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation

论文配图:The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
图 1 · 摘自论文原文
  • 基于维基数据生成带指定流行度的实体数据集
  • 3000个实体在河流/灾害/车款领域,覆盖高低流行度
  • 支持中英文长文本事实性评估,适合模型评测者使用

我们提出一个可配置的流水线,利用维基百科和维基数据生成具有特定特征(如领域、地理位置、流行度)的多语言实体数据集,用于评估大语言模型在长文本生成中的事实性,弥补短文本问答类评测的不足。以RiDiC数据集为例,该数据集包含3000个实体,涵盖河流、自然灾害和汽车型号三个领域,覆盖不同流行度层级。每个实体均配有地理坐标、英中文名称(如有)及对应的英中文维基内容,用于评估模型输出。通过三款大模型在中英文下的生成结果,经第三方事实核查工具检测,发现即使前沿模型在该数据集上仍存在严重幻觉。相关代码、数据及评测脚本已开源,便于多语言长文本事实性评估。

原文摘要 · Abstract (English)

We present a configurable pipeline for generating multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data from Wikipedia and Wikidata. These datasets are intended for evaluating the factuality of LLMs' long-form generation, thereby complementing evaluation based on short-form QA datasets. We present the RiDiC dataset as an example of this approach. RiDiC contains 3,000 entities from three domains -- rivers, natural disasters, and car models -- spanning different popularity tiers. Each entity is accompanied by its geographical location, English and Chinese names (if available) and relevant English and Chinese Wikipedia content, which is used to evaluate LLMs' responses. Generations about RiDiC entities were obtained from three LLMs in English and Chinese. These were then evaluated using a third-party factuality checker, which showed that entities from our dataset caused even frontier models to hallucinate. To facilitate the evaluation of LLMs' long-form factuality in multiple languages, the code, data, and generation/evaluation scripts have been released.

事实性评估长文本生成多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。