评测大模型识别文化错误的能力,发现其准确率仅52%。
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

- 构建多语言文化错误数据集,标注7470个错误片段。
- 最强模型检测错误的F1仅为0.52,难识别深层文化偏差。
- 适合关注AI跨文化适配与评估的研究者参考。
随着大语言模型(LLMs)在全球范围内的部署日益增多,它们被用于各种跨文化场景,如撰写个人信件或构思创意。这些任务本质上具有文化属性:需要语境恰当性、符号共鸣和隐含的文化期待,而母语者能本能感知。因此,一个内容看似合理但对本地读者明显错误的回应仍属不当。现有文化评估基准将文化视为扁平化的事实集合,通过事实验证或规范蕴含方法进行评判,并未检验大模型作为评判者是否能捕捉此类深层文化错误。为此,我们提出JuICE(LLM-Judge识别文化错误的基准),一个包含7,470个细粒度标注的文化与语言错误的多语言数据集,涵盖来自美国、韩国、印度尼西亚和孟加拉国的1,050个问答对,覆盖英语及各国主要语言。基于JuICE,我们发现即使是最强的LLM-judge在错误片段检测任务中也仅达到F1=0.52,且普遍遗漏本地居民轻易识别的深层文化错误。研究结果表明,稳健的文化评估必须超越表层检测,转向考量文化意义的深度与情境性。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural contexts, from drafting personal communications to brainstorming creative ideas. These tasks are inherently cultural: they require contextual appropriateness, symbolic resonance, and tacit cultural expectations that native speakers draw on instinctively, meaning that a response can be factually plausible yet unmistakably wrong to a local reader. Existing cultural benchmarks have treated culture as a flat set of facts via fact verification or norm entailment methods, and have adopted LLM-as-a-Judge without examining whether they can capture such thick cultural errors. To address this gap, we present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh), in both English and their countries' main languages. Using JuICE, we find that even the strongest LLM-judge achieves only an F1 of 0.52 in the erroneous span detection task. Furthermore, LLM-judges consistently miss thick cultural errors that local residents readily identify. Our findings suggest that robust cultural evaluation must move beyond surface-level detection toward frameworks that account for the depth and situatedness of cultural meaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。