提出系统性文化评估方法,让AI评价不再忽视隐含文化偏见。
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
- 从静态事实转向考察评估中隐藏的文化假设。
- 强调研究者自身立场对公平评估的关键影响。
- 适合关注AI公平性与跨文化适配的研究者。
当前主流的大型语言模型文化对齐评估多以‘知识题库’为中心,将文化简化为孤立的事实或价值观,通过选择题或简答题测试,忽略了文化的多元性与互动性。此类方法未意识到文化假设可能渗透于看似‘中立’的评估场景中。本文呼吁推行‘有意文化评估’:系统检视评估全过程中的文化因素,而不仅限于显性文化任务。我们系统分析了文化相关考量在评估中出现的性质、方式及情境,并强调研究者身份立场对实现包容性、文化对齐的NLP研究的重要性。最后讨论了超越现有基准的启示与未来方向,包括发现未知应用,并通过受人机交互启发的参与式方法让社区共同设计评估体系。
原文摘要 · Abstract (English)
The prevailing ``trivia-centered paradigm'' for evaluating the cultural alignment of large language models (LLMs) is increasingly inadequate as these models become more advanced and widely deployed. Existing approaches typically reduce culture to static facts or values, testing models via multiple-choice or short-answer questions that treat culture as isolated trivia. Such methods neglect the pluralistic and interactive realities of culture, and overlook how cultural assumptions permeate even ostensibly ``neutral'' evaluation settings. In this position paper, we argue for \textbf{intentionally cultural evaluation}: an approach that systematically examines the cultural assumptions embedded in all aspects of evaluation, not just in explicitly cultural tasks. We systematically characterize the what, how, and circumstances by which culturally contingent considerations arise in evaluation, and emphasize the importance of researcher positionality for fostering inclusive, culturally aligned NLP research. Finally, we discuss implications and future directions for moving beyond current benchmarking practices, discovering important applications that we don't know exist, and involving communities in evaluation design through HCI-inspired participatory methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。