评测大模型对隐性文化价值观的理解能力,发现小模型微调后可超越大模型。
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
- 构建对话式基准CQBench,评估模型从日常对话中推断隐性文化价值的能力。
- 前沿模型在价值识别上接近人类水平(F1=0.809),但在态度判断上仍有差距(F1=0.622)。
- 仅用500个文化样本微调小模型,性能提升超10%,适合跨文化应用研究者参考。
文化智能(CQ)指理解陌生文化背景的能力,是大语言模型有效服务全球用户的关键。现有研究多关注显性文化规范,忽视日常对话中的隐性价值观。为此,我们提出CQBench,一个基于多角色对话故事的基准,涵盖世界价值观调查与GlobalOpinions中的伦理、宗教、社会等主题。通过自动化数据构建流程,结合一致性、隐含性等验证,最终达成94.5%的人类模型一致性。我们设计三类递增复杂度任务:态度检测、价值选择、价值提取,评估模型在自然对话中识别隐含价值的能力。结果表明,前沿模型如o1在价值选择上达到人类水平(F1=0.809),但在态度检测上仍落后(F1=0.622)。有趣的是,仅用500个文化丰富样本微调小型模型LLaMA-3.2-3B,性能提升超10%,甚至优于o3-mini。本研究揭示了当前大模型文化推理的挑战,并为提升跨文化能力提供可行路径。
原文摘要 · Abstract (English)
Cultural Intelligence (CQ) refers to the ability to understand unfamiliar cultural contexts, a crucial skill for large language models (LLMs) to effectively engage with globally diverse users. Existing studies often focus on explicitly stated cultural norms, but fail to capture the subtle, implicit values that are common in daily conversation. To address this gap, we introduce CQBench, a benchmark specifically designed to assess LLMs' capability to infer implicit cultural values from natural conversational contexts. CQBench consists of multi character conversation based stories using values from the World Value Survey and the GlobalOpinions, with topics including ethical, religious, social, etc. Our automatic dataset construction pipeline integrates rigorous validation procedures (incorporation, consistency, and implicitness checks), achieving a 94.5% human model agreement in the final validation. To leverage CQBench data, we design three tasks of increasing complexity: attitude detection, value selection, and value extraction. These tasks evaluate whether models can detect attitude and recognize values embedded within natural dialogues rather than relying on explicit cultural knowledge. We find that while frontier models like o1 reach human level performance in value selection (0.809 F1), they still fall short in nuanced attitude detection (0.622 F1). Notably, finetuning a smaller LLaMA-3.2-3B on only 500 culturally rich examples improves performance by over 10%, even outperforming o3-mini in some cases. Using CQ-Bench, we provide insights into the current challenges in LLMs' CQ research and suggest practical pathways for enhancing LLMs' cross-cultural reasoning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。