对比多种跨语言一致性提升方法,发现训练后对齐更可靠。
A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

- 统一评测三类模型在三个数据集上的跨语言一致性增强方法。
- 训练后直接分布对齐在所有组合中稳定提升一致性,其他方法依赖回答格式和语言覆盖范围。
- 提升一致性不损害文化敏感问题的差异响应能力,但开放生成可能降低非英语准确率。
多语言大模型常在不同语言中对语义等价问题给出不一致的回答,促使了跨语言一致性(CLC)增强方法的发展。然而现有方法多使用不同模型、任务和评估协议,难以比较其优劣。本文对代表性CLC增强方法进行了统一评估,涵盖推理时干预与训练后改进方法,覆盖三种模型族和三个封闭式问答基准。结果表明,训练后方法普遍更可靠,直接分布对齐在所有模型-数据集组合中持续提升一致性;而其他方法对答案格式和语言覆盖范围更敏感。值得注意的是,跨领域迁移受限,除非源任务与目标任务具有相似输出格式。我们进一步检验了这些方法是否削弱模型在需要时做出文化差异回应的能力。在两个包含文化多样性问题的数据集上,封闭式评估未发现系统性退化;但在开放式生成中,非英语回答准确率偶尔下降。研究强调需同时评估方法在跨域鲁棒性和文化适切性差异上的表现,为后续训练策略与基准建设提供指导。
原文摘要 · Abstract (English)
Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models' ability to respond differently *when needed*, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。