arXiv:2501.18457cs.CL2025-01NAACL被引 17

让大模型跨语言答题更一致,通过自一致性筛选优化对齐

CALM: Unleashing the Cross-Lingual Self-Aligning Ability of Language Model Question Answering

  • 用多语言回答的自一致性筛选目标,负样本来自其他语言答案
  • 在MEDQA和X-CSQA上提升跨语言问答准确率,零样本与检索增强均有效
  • 语言越多效果越好,适合多语言知识对齐研究者使用

大型语言模型在海量多语言语料上预训练,本应实现跨语言文化无关问题的一致回答,但实际表现差异显著。为此,我们探索了语言模型的跨语言自对齐能力(CALM)。针对同一问题,从不同语言中采样多个回答,选取最自一致的回答作为目标,其余作为负样本,采用直接偏好优化(DPO)对齐模型跨语言知识。在MEDQA和X-CSQA数据集上的评估表明,该方法在零样本及检索增强设置下均有效提升跨语言知识问答性能。同时发现,参与训练的语言种类越多,模型准确率与一致性越高。我们还定性分析了跨语言一致性如何促进知识对齐,并探讨了方法的泛化能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are pretrained on extensive multilingual corpora to acquire both language-specific cultural knowledge and general knowledge. Ideally, while LLMs should provide consistent responses to culture-independent questions across languages, we observe significant performance disparities. To address this, we explore the Cross-Lingual Self-Aligning ability of Language Models (CALM) to align knowledge across languages. Specifically, for a given question, we sample multiple responses across different languages and select the most self-consistent response as the target, leaving the remaining responses as negative examples. We then employ direct preference optimization (DPO) to align the model's knowledge across different languages. Evaluations on the MEDQA and X-CSQA datasets demonstrate CALM's effectiveness in enhancing cross-lingual knowledge question answering, both in zero-shot and retrieval-augmented settings. We also found that increasing the number of languages involved in CALM training leads to higher accuracy and consistency. We offer a qualitative analysis of how cross-lingual consistency can enhance knowledge alignment and explore the method's generalizability.

跨语言知识对齐大模型问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。