arXiv:2503.21813cs.AIcs.CL2025-03被引 1

构建首个面向本体匹配的LLM幻觉基准数据集,助力识别与抑制幻觉。

OAEI-LLM-T: A TBox Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching

  • 基于OAEI的7个TBox数据集,生成10种LLM在本体匹配中的幻觉数据
  • 归纳出两类六小类幻觉模式,覆盖常见错误类型
  • 可用于评估和优化用于本体匹配的LLM,适合知识图谱研究者

大型语言模型(LLM)在下游任务中常出现幻觉。为应对基于LLM的本体匹配(OM)系统中幻觉这一重大挑战,我们提出新基准数据集OAEI-LLM-T。该数据集源自Ontology Alignment Evaluation Initiative(OAEI)的七个TBox数据集,记录了十种不同LLM在执行OM任务时产生的幻觉。这些针对本体匹配的幻觉被分为两大类六小类。我们展示了该数据集在构建OM任务的LLM排行榜以及微调用于OM的LLM方面的实用性。

原文摘要 · Abstract (English)

Hallucinations are often inevitable in downstream tasks using large language models (LLMs). To tackle the substantial challenge of addressing hallucinations for LLM-based ontology matching (OM) systems, we introduce a new benchmark dataset OAEI-LLM-T. The dataset evolves from seven TBox datasets in the Ontology Alignment Evaluation Initiative (OAEI), capturing hallucinations of ten different LLMs performing OM tasks. These OM-specific hallucinations are organised into two primary categories and six sub-categories. We showcase the usefulness of the dataset in constructing an LLM leaderboard for OM tasks and for fine-tuning LLMs used in OM tasks.

本体匹配幻觉检测大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。