arXiv:2605.29897cs.CL2026-05

首个可解释的文化偏差检测工具,自动识别并分析大模型输出中的文化错误。

ExCAM: Explainable Cultural Awareness Metrics

论文配图:ExCAM: Explainable Cultural Awareness Metrics
图 1 · 摘自论文原文
  • 基于合成错误重构九个基准数据集,构建可解释的文化评估指标。
  • 在平衡测试集上误差检测准确率达80%,优于GPT-5等基线模型。
  • 适合研究文化偏见、提升生成内容公平性的AI开发者与伦理研究人员。

评估大语言模型的文化意识对确保生成文本的公平性及全球应用的泛化能力至关重要。现有基准通过问答或文本生成任务考察饮食、压力情境下的行为等文化要素,但构建需耗时且昂贵的人工标注。此外,针对自由文本的文化意识评估工具稀缺,且多依赖过时方法。为此,我们提出ExCAM——首个专门用于识别、评分并解释指令-输出对中文化错误的可解释评估指标。为训练与评估ExCAM,我们构建了ExCAM40k数据集,整合九个现有基准,并通过合成错误增强。相比多个基线(包括GPT-5),ExCAM在平衡测试集上实现最高达80%的错误检测准确率。这为自由文本层面细粒度、可解释的文化评估开辟了新路径。

原文摘要 · Abstract (English)

Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent benchmarks explore cultural goods like food or values like behavior in stressful situations through the lens of question answering or text generation tasks. However, creating these benchmarks requires time-intensive and costly human annotations. Also, benchmarks that evaluate cultural awareness in free text are scarce and often rely on dated evaluation mechanisms. To address this gap, we introduce ExCAM, an Explainable Cultural Awareness Metric, which is, to our knowledge, the first dedicated evaluation metric that identifies, rates and explains cultural errors in instruction-output pairs. To train and evaluate ExCAM, we introduce ExCAM40k, a dataset comprised of nine existing benchmarks that we reformat and enhance with synthetic errors. Compared to several baselines, including GPT-5, ExCAM achieves the highest error detection rate with up to 80% accuracy on a balanced test set. Therefore, ExCAM opens the pathway towards fine-grained and explainable cultural evaluation of free text.

文化偏见可解释性评估指标LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。