arXiv:2409.05806cs.CLcs.AI2024-09ACL被引 4

首个中文知识纠错数据集,专攻语言、事实与逻辑错误

CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs

  • 构建七类中文知识样本,覆盖古诗成语与网络语料
  • 揭示大模型在汉语文化结构理解上的明显短板
  • 适合研究中文LLM纠错与文化知识建模的学者使用

中文作为语义丰富、结构复杂的语言系统,包含古诗词、谚语、成语等独特文化元素。然而当前大语言模型在这些特定领域存在明显局限,亟需综合性数据集以评估、持续更新并优化其文化语境下的语言能力。为此,我们提出CKnowEdit,首个专注于纠正中文语言、事实与逻辑错误的知识编辑数据集。数据涵盖七类知识,来源包括古典文献、成语典故及百度贴吧等网络语料,充分考虑汉语特有的复调性、对仗性与逻辑结构。通过分析该数据集,我们揭示了现有大模型在掌握汉语深层特征方面的挑战。此外,对前沿知识编辑技术的评估表明,中文知识修正仍有较大提升空间。代码与数据集已开源。

原文摘要 · Abstract (English)

Chinese, as a linguistic system rich in depth and complexity, is characterized by distinctive elements such as ancient poetry, proverbs, idioms, and other cultural constructs. However, current Large Language Models (LLMs) face limitations in these specialized domains, highlighting the need for the development of comprehensive datasets that can assess, continuously update, and progressively improve these culturally-grounded linguistic competencies through targeted training optimizations. To address this gap, we introduce CKnowEdit, the first-ever Chinese knowledge editing dataset designed to correct linguistic, factual, and logical errors in LLMs. We collect seven types of knowledge from a wide range of sources, including classical texts, idioms, and content from Baidu Tieba Ruozhiba, taking into account the unique polyphony, antithesis, and logical structures inherent in the Chinese language. By analyzing this dataset, we highlight the challenges current LLMs face in mastering Chinese. Furthermore, our evaluation of state-of-the-art knowledge editing techniques reveals opportunities to advance the correction of Chinese knowledge. Code and dataset are available at https://github.com/zjunlp/EasyEdit.

知识编辑中文LLM文化认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。