arXiv:2602.09444cs.CLcs.AI2026-02中稿 · the First Workshop…

提出新指标衡量句子文化特异性,可精准区分特定文化与通用表达

Conceptual Cultural Index: A Metric for Cultural Specificity via Relative Generality

  • 通过目标文化内与跨文化泛化度对比计算句子文化特异性
  • 在400句测试中,文化特异句得分更高,通用句得分更低
  • 相比直接模型评分,分类准确率提升超10个百分点

大型语言模型在多文化场景中应用日益广泛,但句子层面的文化特异性系统评估仍不充分。本文提出概念文化指数(CCI),用于估计句子级别的文化特异性。CCI定义为在目标文化内的泛化度估计值与其它文化平均泛化度估计值之差。该设计使用户可通过比较设置灵活控制文化范围,并具备可解释性,因分数源自底层泛化度估计。我们在400个句子(200个文化特异句,200个通用句)上验证了CCI,其得分分布符合预期:文化特异句得分更高,通用句得分更低。在二元可分性任务中,CCI优于直接使用LLM打分,对目标文化专精的模型,AUC提升超过10点。代码已开源于https://github.com/IyatomiLab/CCI。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in multicultural settings; however, systematic evaluation of cultural specificity at the sentence level remains underexplored. We propose the Conceptual Cultural Index (CCI), which estimates cultural specificity at the sentence level. CCI is defined as the difference between the generality estimate within the target culture and the average generality estimate across other cultures. This formulation enables users to operationally control the scope of culture via comparison settings and provides interpretability, since the score derives from the underlying generality estimates. We validate CCI on 400 sentences (200 culture-specific and 200 general), and the resulting score distribution exhibits the anticipated pattern: higher for culture-specific sentences and lower for general ones. For binary separability, CCI outperforms direct LLM scoring, yielding more than a 10-point improvement in AUC for models specialized to the target culture. Our code is available at https://github.com/IyatomiLab/CCI .

文化特异性语言模型评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。