arXiv:2607.20410cs.CL2026-07

首个面向斯里兰卡文化价值观的对齐资源,提升多语言模型本土适配性。

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

论文配图:LKValues: Aligning Large Language Models with Sri Lankan Societal Values
图 1 · 摘自论文原文
  • 基于205人三语调查构建40项本地主流价值观
  • 创建含15万条场景数据的僧伽罗语-英语指令集和千例评测基准
  • 验证大模型在僧伽罗语中仍存文化对齐缺陷,微调后显著改善输出质量

大型语言模型的价值对齐常偏向西方规范,导致在斯里兰卡等多语言社会中忽视本地文化。现有基准未涵盖僧伽罗语语境下的价值观,阻碍了文化敏感的评估与微调。为此,我们提出LKValues,首个基于实地调研的斯里兰卡价值对齐资源。通过对205名受访者进行三语调查,结合全球框架与大模型生成的本地构念,提炼出40项多数认可的社会价值观。基于这些价值观,构建了包含15万条情景式实例的僧伽罗语-英语指令语料库LKvaluesIT,以及包含1000个实例的价值敏感评测基准LKvaluesBench。我们评估了多个专有及开源大模型,并对三个开源基础模型(Qwen3.5-4B-Base、Qwen3.5-9B-Base、Aya-Expanse-8B-Base)进行微调。实验表明,即使新模型仍存在低资源与文化对齐差距。使用LKValues微调可提升Qwen系列模型在英、僧两语的表现,减少无效输出与跨语言差异,但收益取决于模型家族。这证明了LKValues在嵌入斯里兰卡价值观方面的有效性,提供了一套可复用的低资源国家级多元价值观对齐方法。数据集已公开于https://github.com/NextME14/LKValues。

原文摘要 · Abstract (English)

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning. To bridge this gap, we propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment. From a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs, we derive 40 majority-endorsed societal values. Using these values, we construct LKvaluesIT, a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances, and LKvaluesBench, a value-sensitive evaluation benchmark of 1,000 instances. We evaluate a set of proprietary and open-weight LLMs with LKvaluesBench. We fine-tune three open-weight base models (Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base). Our experiments show that newer and larger LLMs still exhibit low-resource and cultural value-alignment gaps. LKValues fine-tuning improves Qwen-family models in English and Sinhala, reducing invalid outputs and cross-lingual disparities, though gains remain model-family dependent. These highlight LKValues efficacy in embedding Sri Lankan values, offering a replicable pipeline for low-resource, country-specific pluralist value alignment. The dataset is publicly available at https://github.com/NextME14/LKValues.

价值观对齐多语言文化适配斯里兰卡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。