arXiv:2410.10489cs.CLcs.AI2024-10被引 6

低资源语言的数字数据量决定大模型对社会价值观的反映能力。

Cultural Fidelity in Large-Language Models: An Evaluation of Online Language Resources as a Driver of Model Performance in Value Representation

  • 用21个语种国家的数据验证语言数字资源与模型价值表征能力的关系。
  • 高资源语言下GPT-4o误差率仅为低资源语言的1/5,相关性达44%。
  • 适合关注AI公平性、多语言模型优化的研究者和政策制定者。

大型语言模型的训练数据蕴含社会价值观,使其对特定语言文化的熟悉度随之提升。我们的分析发现,GPT-4o在反映各国社会价值观方面的能力(以世界价值观调查为指标)中,有44%的方差可由该语言的数字资源可用性解释。值得注意的是,资源最少的语言误差率比资源最丰富的语言高出五倍以上。对于GPT-4-turbo,这一相关性上升至72%,表明其在非英语语言上的文化熟悉度已超越仅依赖网络抓取数据的努力。本研究构建了该领域最大且最稳健的数据集之一,包含21个国别-语言组合,每组涵盖94个经母语者验证的调查问题。结果凸显了大模型表现与目标语言数字数据可得性的紧密关联。低资源语言(尤其是全球南方)表现较弱可能加剧数字鸿沟。我们讨论了应对策略,包括从零开始构建多语言模型,以及通过多样化语言数据集进行微调,如非洲语言倡议所示。

原文摘要 · Abstract (English)

The training data for LLMs embeds societal values, increasing their familiarity with the language's culture. Our analysis found that 44% of the variance in the ability of GPT-4o to reflect the societal values of a country, as measured by the World Values Survey, correlates with the availability of digital resources in that language. Notably, the error rate was more than five times higher for the languages of the lowest resource compared to the languages of the highest resource. For GPT-4-turbo, this correlation rose to 72%, suggesting efforts to improve the familiarity with the non-English language beyond the web-scraped data. Our study developed one of the largest and most robust datasets in this topic area with 21 country-language pairs, each of which contain 94 survey questions verified by native speakers. Our results highlight the link between LLM performance and digital data availability in target languages. Weaker performance in low-resource languages, especially prominent in the Global South, may worsen digital divides. We discuss strategies proposed to address this, including developing multilingual LLMs from the ground up and enhancing fine-tuning on diverse linguistic datasets, as seen in African language initiatives.

大模型语言公平数字鸿沟价值观建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。