跨语言道德推理测试揭示低资源语言关键作用
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
- 构建多语言道德推理基准,覆盖五种语言与三重上下文复杂度
- 低资源语言在长文本中表现显著下降,越南语降幅最明显
- 微调后低资源语言对多语言推理提升更显著,凸显其重要性
本文提出多语言道德推理基准(MMRB),评估大语言模型在五种类型差异大的语言中,于句子、段落和文档三种上下文复杂度下的道德推理能力。结果表明,随着上下文复杂度上升,模型性能普遍下降,尤其在低资源语言如越南语中更为明显。通过使用精选的单语数据对开源模型LLaMA-3-8B进行对齐与污染微调,发现低资源语言对多语言推理的影响反而强于高资源语言,凸显其在多语言自然语言处理中的关键作用。
原文摘要 · Abstract (English)
In this paper, we introduce the Multilingual Moral Reasoning Benchmark (MMRB) to evaluate the moral reasoning abilities of large language models (LLMs) across five typologically diverse languages and three levels of contextual complexity: sentence, paragraph, and document. Our results show moral reasoning performance degrades with increasing context complexity, particularly for low-resource languages such as Vietnamese. We further fine-tune the open-source LLaMA-3-8B model using curated monolingual data for alignment and poisoning. Surprisingly, low-resource languages have a stronger impact on multilingual reasoning than high-resource ones, highlighting their critical role in multilingual NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。