arXiv:2601.17397cs.CL2026-01

构建中文原生知识编辑基准,发现中英知识修改互不传递。

CLM-Bench: Benchmarking and Analyzing Cross-lingual Misalignment of LLMs in Knowledge Editing

  • 以中文文化为本构建1010组反事实数据对,避免翻译偏差。
  • 实测多语言模型在中英间编辑无法互相迁移,跨语言错位严重。
  • 适合关注多语言模型知识更新与文化适应性的研究者。

知识编辑(KE)已成为无需重新训练即可更新大语言模型(LLM)事实信息的有前景范式。然而,多语言知识编辑(MKE)的发展目前受限于存在偏见的评估框架。我们观察到,现有MKE基准通常通过机械翻译英文中心数据集得到目标语言版本(如英译中),这引入了翻译误差,并忽略了目标语言特有的文化实体,未能反映LLM的真实知识分布。为此,我们提出CLM-Bench,一个基于中文原生方法构建的文化感知基准。我们整理了1,010组根植于中国文化的高质量CounterFact对,并与英文对应项对齐。利用该基准,我们在代表性模型(如Llama-3、Qwen2)上进行广泛实验,揭示出显著的跨语言错位:一种语言中的编辑独立运行,无法传播至另一种语言。通过逐层表征分析,我们提供几何解释:中英文编辑向量几乎正交——位于分离子空间中;而混合语言编辑则表现出这些向量的线性叠加性。研究结果挑战了当前跨语言迁移方法的有效性,强调了文化原生基准的重要性。

原文摘要 · Abstract (English)

Knowledge Editing (KE) has emerged as a promising paradigm for updating facts in Large Language Models (LLMs) without retraining. However, progress in Multilingual Knowledge Editing (MKE) is currently hindered by biased evaluation frameworks. We observe that existing MKE benchmarks are typically constructed by mechanically translating English-centric datasets into target languages (e.g., English-to-Chinese). This approach introduces translation artifacts and neglects culturally specific entities native to the target language, failing to reflect the true knowledge distribution of LLMs. To address this, we propose CLM-Bench, a culture-aware benchmark constructed using a native Chinese-first methodology. We curate 1,010 high-quality CounterFact pairs rooted in Chinese cultural contexts and align them with English counterparts. Using CLM-Bench, we conduct extensive experiments on representative LLMs (e.g., Llama-3, Qwen2) and reveal a significant Cross-lingual Misalignment: edits in one language function independently and fail to propagate to the other. We further provide a geometric explanation via layer-wise representation analysis, demonstrating that edit vectors for Chinese and English are nearly orthogonal -- residing in disjoint subspaces -- while mixed-lingual editing exhibits linear additivity of these vectors. Our findings challenge the effectiveness of current methods in cross-lingual transfer and underscore the importance of culturally native benchmarks.

知识编辑多语言模型文化对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。