arXiv:2511.12014cs.CLcs.HC2025-11被引 5

用真实情境评估大模型文化理解力,比传统方法更准更稳。

CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs

  • 设计真实场景测试模型的文化推理能力
  • 新指标揭示主流模型文化理解深度不足
  • 适合研究文化对齐与评估的学者使用

大型语言模型在多元文化环境中应用日益广泛,但现有文化能力评估仍显不足。现有方法多聚焦于脱离语境的正确性或强制选择判断,忽视了恰当回应所需的文化理解与推理。为填补这一空白,我们提出一组基准测试,不直接探测抽象规范或孤立陈述,而是呈现需要基于文化背景进行推理的真实情境。除标准的Exact Match指标外,还引入四种互补指标(覆盖率、特异性、情感色彩、连贯性)以捕捉响应质量的不同维度。对前沿模型的实证分析表明,浅层评估系统性高估文化能力,且结果波动大;而深层评估能揭示推理深度差异,降低方差,提供更稳定、可解释的文化理解信号。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in culturally diverse environments, yet existing evaluations of cultural competence remain limited. Existing methods focus on de-contextualized correctness or forced-choice judgments, overlooking the need for cultural understanding and reasoning required for appropriate responses. To address this gap, we introduce a set of benchmarks that, instead of directly probing abstract norms or isolated statements, present models with realistic situational contexts that require culturally grounded reasoning. In addition to the standard Exact Match metric, we introduce four complementary metrics (Coverage, Specificity, Connotation, and Coherence) to capture different dimensions of model's response quality. Empirical analysis across frontier models reveals that thin evaluation systematically overestimates cultural competence and produces unstable assessments with high variance. In contrast, thick evaluation exposes differences in reasoning depth, reduces variance, and provides more stable, interpretable signals of cultural understanding.

文化对齐评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。