arXiv:2501.01056cs.CLcs.AI2025-01被引 13

警告大模型可能造成文化消亡,尤其对数字资源少的群体。

Risks of Cultural Erasure in Large Language Models

  • 区分文化遗漏与简化,评估模型如何呈现全球文化
  • 发现模型在描述地点和旅行推荐中常弱化少数文化复杂性
  • 呼吁建立可量化的评估标准,纳入历史权力不平等考量

大型语言模型正深度融入影响社会知识生产与传播的应用,如搜索、在线教育和旅行规划。这使得模型将塑造人们对全球文化的认知与互动方式,因而必须关注其训练数据中哪些知识体系与视角被体现。尽管已有研究关注全球文化表征分布的缺口,但仍缺乏针对语言模型跨文化影响的细粒度评估基准。本文主张发展可量化的评价体系,以揭示历史权力不均对文化表征差异的影响,特别关注数字语料中已处于弱势的文化。我们分析两种文化消亡形式:一是‘遗漏’——文化完全未被提及;二是‘简化’——用单一维度呈现丰富文化。通过考察模型对世界各地的描述及旅行推荐中的文化表征,发现当前模型显著削弱了边缘文化的复杂性。本文提出,应将这些社会文化考量转化为可操作的标准评估与基准,供自然语言处理社区与应用开发者参考。

原文摘要 · Abstract (English)

Large language models are increasingly being integrated into applications that shape the production and discovery of societal knowledge such as search, online education, and travel planning. As a result, language models will shape how people learn about, perceive and interact with global cultures making it important to consider whose knowledge systems and perspectives are represented in models. Recognizing this importance, increasingly work in Machine Learning and NLP has focused on evaluating gaps in global cultural representational distribution within outputs. However, more work is needed on developing benchmarks for cross-cultural impacts of language models that stem from a nuanced sociologically-aware conceptualization of cultural impact or harm. We join this line of work arguing for the need of metricizable evaluations of language technologies that interrogate and account for historical power inequities and differential impacts of representation on global cultures, particularly for cultures already under-represented in the digital corpora. We look at two concepts of erasure: omission: where cultures are not represented at all and simplification i.e. when cultural complexity is erased by presenting one-dimensional views of a rich culture. The former focuses on whether something is represented, and the latter on how it is represented. We focus our analysis on two task contexts with the potential to influence global cultural production. First, we probe representations that a language model produces about different places around the world when asked to describe these contexts. Second, we analyze the cultures represented in the travel recommendations produced by a set of language model applications. Our study shows ways in which the NLP community and application developers can begin to operationalize complex socio-cultural considerations into standard evaluations and benchmarks.

文化消亡语言模型评估基准代表性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。