挑战主流文化对齐评估方式,发现开放问答更真实反映大模型文化倾向。
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
- 用开放问答替代封闭选择题,避免强行匹配选项带来的偏差。
- 在非约束场景下,模型文化对齐表现显著提升,响应更自然一致。
- 建议采用具体文化指标评估,适合关注模型跨文化适应的研究者。
大量研究依赖封闭式多选问卷评估大语言模型(LLMs)的文化对齐性。本文挑战这一受限评估范式,探索更贴近现实的开放式方法。以世界价值观调查(WVS)和霍夫斯泰德文化维度为例,我们发现,在无需强制选择的非约束环境中,LLMs展现出更强的文化对齐能力。此外,仅调整选项顺序便导致输出不一致,暴露了封闭式评估的局限性。研究呼吁建立更稳健、灵活的评估框架,聚焦具体文化代理指标,推动对文化对齐更精细、准确的衡量。
原文摘要 · Abstract (English)
A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more realistic, unconstrained approaches. Using the World Values Survey (WVS) and Hofstede Cultural Dimensions as case studies, we demonstrate that LLMs exhibit stronger cultural alignment in less constrained settings, where responses are not forced. Additionally, we show that even minor changes, such as reordering survey choices, lead to inconsistent outputs, exposing the limitations of closed-style evaluations. Our findings advocate for more robust and flexible evaluation frameworks that focus on specific cultural proxies, encouraging more nuanced and accurate assessments of cultural alignment in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。