构建印度文化专有概念数据集,评估大模型在本土文本适配中的文化能力。
DIWALI: Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
- 针对印度17个文化维度、36个次区域构建8000+文化概念数据集
- 发现主流大模型在次区域文化覆盖上存在选择性缺失和浅层适配
- 结合人类与模型评分,提供可复现的文化适应性评估框架
大型语言模型虽广泛应用,但因缺乏文化知识与能力,常产生文化偏差。现有文化专有概念(CSI)数据集多聚焦区域层面,且可能存在误标。为此,本文构建了面向印度文化的新型CSI数据集,涵盖17个文化维度、36个次区域,包含约8000个文化概念。为评估大模型在文化文本适配任务中的文化胜任力,我们基于该数据集,采用大模型作为裁判并结合多元社会人口背景的人类评估进行量化分析。结果表明,所有考察的大模型均存在次区域覆盖不均和表面化适配问题。数据集与代码已开源:https://huggingface.co/datasets/nlip/DIWALI,项目主页:https://nlip-lab.github.io/nlip/publications/diwali/,代码库:https://github.com/pramitsahoo/culture-evaluation。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used in various tasks and applications. However, despite their wide capabilities, they are shown to lack cultural alignment \citep{ryan-etal-2024-unintended, alkhamissi-etal-2024-investigating} and produce biased generations \cite{naous-etal-2024-beer} due to a lack of cultural knowledge and competence. Evaluation of LLMs for cultural awareness and alignment is particularly challenging due to the lack of proper evaluation metrics and unavailability of culturally grounded datasets representing the vast complexity of cultures at the regional and sub-regional levels. Existing datasets for culture specific items (CSIs) focus primarily on concepts at the regional level and may contain false positives. To address this issue, we introduce a novel CSI dataset for Indian culture, belonging to 17 cultural facets. The dataset comprises ~8k cultural concepts from 36 sub-regions. To measure the cultural competence of LLMs on a cultural text adaptation task, we evaluate the adaptations using the CSIs created, LLM as Judge, and human evaluations from diverse socio-demographic region. Furthermore, we perform quantitative analysis demonstrating selective sub-regional coverage and surface-level adaptations across all considered LLMs. Our dataset is available here: https://huggingface.co/datasets/nlip/DIWALI, project webpage https://nlip-lab.github.io/nlip/publications/diwali/, and our codebase with model outputs can be found here: https://github.com/pramitsahoo/culture-evaluation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。