arXiv:2409.14038cs.AIcs.CL2024-09被引 10

构建首个专用于评估大模型在本体匹配中幻觉的基准数据集

OAEI-LLM: A Benchmark Dataset for Understanding Large Language Model Hallucinations in Ontology Matching

  • 扩展OAEI数据集,加入大模型幻觉标注
  • 首次系统评测大模型在本体匹配中的错误生成模式
  • 适合研究大模型可靠性与知识对齐的学者

大语言模型(LLMs)在领域特定下游任务中常出现幻觉,本体匹配(OM)也不例外。随着LLMs被广泛应用于OM,亟需基准数据集来深入理解其幻觉现象。本文提出的OAEI-LLM是基于本体对齐评估倡议(OAEI)数据集的扩展版本,专门用于评估LLM在本体匹配任务中的特异性幻觉。文中详述了数据集构建方法与模式扩展策略,并提供了潜在应用场景示例。

原文摘要 · Abstract (English)

Hallucinations of large language models (LLMs) commonly occur in domain-specific downstream tasks, with no exception in ontology matching (OM). The prevalence of using LLMs for OM raises the need for benchmarks to better understand LLM hallucinations. The OAEI-LLM dataset is an extended version of the Ontology Alignment Evaluation Initiative (OAEI) datasets that evaluate LLM-specific hallucinations in OM tasks. We outline the methodology used in dataset construction and schema extension, and provide examples of potential use cases.

大模型幻觉本体匹配基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。