让NLP真正懂文化:超越语言列表,关注语言使用的生态
Toward Culturally Grounded Natural Language Processing
- 从语言列表转向研究语言的使用生态,包括机构、社区和媒介
- 发现数据覆盖之外,分词、提示语言、评估设计也影响跨语言表现
- 适合关注公平性、文化适配与社区参与的NLP研究者
多语言NLP常被视为实现全球包容的路径,但语言覆盖与文化理解能力往往脱节。本文综合50余篇论文,涵盖多语言性能不平等、跨语言迁移、文化感知评估、文化对齐、多模态基准、基准设计批判及社区驱动的数据实践。研究表明,训练数据覆盖虽关键,但结果还受分词方式、提示语言、翻译后的基准设计、文化导向标注、模态类型以及评估数据的作者与验证者影响。我们主张,文化根基的NLP应跳出将语言视为基准表中孤立条目的模式,转而建模语言的传播生态:即语言在其中被使用的制度、书写系统、领域、模态和社群。为此提出分层评估与报告框架,包含代表性审计、混合式数据采集、生态效度、社区验证、适应性溯源、语言内变异分析,以及活态文化资源的持续维护。
原文摘要 · Abstract (English)
Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper synthesizes over 50 papers spanning multilingual performance inequality, cross-lingual transfer, culture-aware evaluation, cultural alignment, multimodal benchmarks, benchmark-design critique, and community-grounded data practices. Across this literature, training data coverage remains important, but outcomes are also shaped by tokenization, prompt language, translated benchmark design, culturally grounded supervision, modality, and who authors or validates evaluation data. We argue that culturally grounded NLP should move beyond treating languages as isolated rows in benchmark tables and instead model communicative ecologies: the institutions, scripts, domains, modalities, and communities through which language is used. We propose a layered evaluation and reporting agenda centered on representation audits, mixed elicitation, ecological validity, community validation, adaptation provenance, within-language variation, and maintenance of living cultural resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。