arXiv:2511.21322cs.HCcs.AI2025-11被引 10

评估大模型生成故事中的文化误表征,发现88%存在错误。

TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories

  • 基于印度本地人经验构建文化误表征分类体系
  • 88%生成故事含文化错误,低资源语言更严重
  • 提供可复用的文化知识评测题库,适合开发者优化模型

全球数百万用户使用AI聊天机器人满足创作需求,引发对其如何呈现多元文化的关注。然而,开放式任务中文化表征的评估仍具挑战且研究不足。本文提出TALES,针对印度多样文化身份的生成故事进行文化误表征评估。首先,通过9场焦点小组与15份个人问卷,汇聚印度本土人士经验,构建TALES-Tax文化误表征分类体系。基于该体系,对6个模型开展大规模标注研究,涵盖2925条标注,来自印度71个地区、14种语言、具备母语能力的108名标注者。结果显示,88%的生成故事存在文化误表征,且在中低资源语言及印度近郊地区故事中尤为普遍。此外,将标注结果转化为独立的TALES-QA问答库,用于评估模型文化知识水平。

原文摘要 · Abstract (English)

Millions of users across the globe turn to AI chatbots for their creative needs, inviting widespread interest in understanding how they represent diverse cultures. However, evaluating cultural representations in open-ended tasks remains challenging and underexplored. In this work, we present TALES, an evaluation of cultural misrepresentations in LLM-generated stories for diverse Indian cultural identities. First, we develop TALES-Tax, a taxonomy of cultural misrepresentations by collating insights from participants with lived experiences in India through focus groups (N=9) and individual surveys (N=15). Using TALES-Tax, we evaluate 6 models through a large-scale annotation study spanning 2925 annotations from 108 annotators with lived experience and native language proficiency from across 71 regions in India and 14 languages. Concerningly, we find that 88% of the generated stories contain misrepresentations, and such errors are more prevalent in mid- and low-resourced languages and stories based in peri-urban regions in India. We also transform the annotations into TALES-QA, a standalone question bank to evaluate the cultural knowledge of models.

文化表征大模型评估印度研究偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。