arXiv:2505.12345cs.CL2025-05NeurIPS被引 9

构建首个覆盖多领域的大模型知识编辑统一评测基准

UniEdit: A Unified Knowledge Editing Benchmark for Large Language Models

  • 基于25个常见领域的开放知识图谱构建编辑样本
  • 设计多跳邻域采样算法,全面评估编辑的连锁影响
  • 适配多种大模型与编辑方法,助力研究者对比优化

模型编辑旨在通过高效调整内部参数提升大语言模型(LLMs)的准确性与可靠性。当前多数大模型编辑数据集局限于狭窄知识领域,且评估范围有限,常忽略编辑需求的广泛性及修改带来的多样化连锁效应。为此,我们提出UniEdit,一个基于开放领域知识的统一编辑评测基准。首先,从五个主要类别下的25个常见领域中选取实体,利用开放知识图谱中的丰富三元组信息,确保知识覆盖全面。为解决编辑的泛化性与局部性问题,设计了邻域多跳链采样(NMCS)算法,基于给定知识片段采样子图,以涵盖全面的涟漪效应。最后,使用专有大模型将采样的知识子图转化为自然语言文本,保障语法准确性和句法多样性。大量统计分析验证了UniEdit在规模、全面性与多样性上的优势。我们在多个大模型与编辑器上进行了综合实验,分析其在开放知识领域和不同评估标准下的表现,揭示了各类方法的优劣,为未来研究提供重要参考。

原文摘要 · Abstract (English)

Model editing aims to enhance the accuracy and reliability of large language models (LLMs) by efficiently adjusting their internal parameters. Currently, most LLM editing datasets are confined to narrow knowledge domains and cover a limited range of editing evaluation. They often overlook the broad scope of editing demands and the diversity of ripple effects resulting from edits. In this context, we introduce UniEdit, a unified benchmark for LLM editing grounded in open-domain knowledge. First, we construct editing samples by selecting entities from 25 common domains across five major categories, utilizing the extensive triple knowledge available in open-domain knowledge graphs to ensure comprehensive coverage of the knowledge domains. To address the issues of generality and locality in editing, we design an Neighborhood Multi-hop Chain Sampling (NMCS) algorithm to sample subgraphs based on a given knowledge piece to entail comprehensive ripple effects to evaluate. Finally, we employ proprietary LLMs to convert the sampled knowledge subgraphs into natural language text, guaranteeing grammatical accuracy and syntactical diversity. Extensive statistical analysis confirms the scale, comprehensiveness, and diversity of our UniEdit benchmark. We conduct comprehensive experiments across multiple LLMs and editors, analyzing their performance to highlight strengths and weaknesses in editing across open knowledge domains and various evaluation criteria, thereby offering valuable insights for future research endeavors.

大模型编辑知识图谱评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。