让大模型真正理解文化,而非简单套用标签
'Too much alignment; not enough culture': Re-balancing cultural alignment practices in LLMs
- 引入人类学的深度描述法,要求模型输出体现文化深层含义
- 提出三个文化对齐必要条件:范围适配、表达细腻、上下文锚定
- 倡导跨学科合作,用民族志方法评估模型文化敏感性
尽管文化对齐在人工智能研究中日益重要,但现有方法主要依赖量化基准和简化代理,难以捕捉人类文化的深层复杂性和情境依赖性。当前对齐实践常将文化简化为静态人口分类或表面文化事实,忽略了何为真正文化对齐的核心问题。本文主张将社会科学研究中的诠释性定性方法融入大语言模型(LLMs)对齐实践,借鉴克利福德·格尔茨的“厚描述”概念,提出AI系统应生成反映深层文化意义的“厚输出”,并基于用户提供的上下文与意图。我们明确了三种成功文化对齐的必要条件:文化表征范围适度、输出具备细腻性、输出锚定于提示中隐含的文化语境。最后,呼吁跨学科协作及采用定性、民族志式评估方法,推动开发真正具备文化敏感性、伦理责任与人类复杂性的AI系统。
原文摘要 · Abstract (English)
While cultural alignment has increasingly become a focal point within AI research, current approaches relying predominantly on quantitative benchmarks and simplistic proxies fail to capture the deeply nuanced and context-dependent nature of human cultures. Existing alignment practices typically reduce culture to static demographic categories or superficial cultural facts, thereby sidestepping critical questions about what it truly means to be culturally aligned. This paper argues for a fundamental shift towards integrating interpretive qualitative approaches drawn from social sciences into AI alignment practices, specifically in the context of Large Language Models (LLMs). Drawing inspiration from Clifford Geertz's concept of "thick description," we propose that AI systems must produce outputs that reflect deeper cultural meanings--what we term "thick outputs"-grounded firmly in user-provided context and intent. We outline three necessary conditions for successful cultural alignment: sufficiently scoped cultural representations, the capacity for nuanced outputs, and the anchoring of outputs in the cultural contexts implied within prompts. Finally, we call for cross-disciplinary collaboration and the adoption of qualitative, ethnographic evaluation methods as vital steps toward developing AI systems that are genuinely culturally sensitive, ethically responsible, and reflective of human complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。