arXiv:2507.09629cs.CL2025-07被引 1

首次研究阿拉伯语知识编辑,发现指令微调方法更稳健。

An Exploration of Knowledge Editing for Arabic

  • 对比四种编辑方法在阿拉伯语中的表现,评估跨语言迁移能力。
  • 参数方法在跨语言场景下效果差,指令微调模型更稳定。
  • 提出多语言联合训练的LTE,提升阿拉伯语编辑效果与迁移性。

尽管知识编辑(KE)在英语中已有广泛研究,但在形态丰富的语言如阿拉伯语中的表现仍缺乏探讨。本文首次系统研究阿拉伯语的知识编辑问题,评估了四种方法(ROME、MEMIT、ICE、LTE)在阿拉伯语版ZsRE和Counterfact基准上的表现,涵盖多语言与跨语言设置。实验基于Llama-2-7B-chat模型发现,参数型方法在跨语言迁移中表现不佳,而指令微调方法更具鲁棒性。本文将学习编辑(LTE)扩展至多语言场景,通过阿拉伯语-英语联合训练,显著提升编辑能力和迁移性能。研究释放了阿拉伯语知识编辑基准及LTE多语言训练数据,以支持后续研究。

原文摘要 · Abstract (English)

While Knowledge Editing (KE) has been widely explored in English, its behavior in morphologically rich languages like Arabic remains underexamined. In this work, we present the first study of Arabic KE. We evaluate four methods (ROME, MEMIT, ICE, and LTE) on Arabic translations of the ZsRE and Counterfact benchmarks, analyzing both multilingual and cross-lingual settings. Our experiments on Llama-2-7B-chat show that parameter-based methods struggle with cross-lingual generalization, while instruction-tuned methods perform more robustly. We extend Learning-To-Edit (LTE) to a multilingual setting and show that joint Arabic-English training improves both editability and transfer. We release Arabic KE benchmarks and multilingual training for LTE data to support future research.

知识编辑阿拉伯语多语言指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。