用社区视角重构AI文化表达评估,避免刻板化与抽象化
The Case for "Thick Evaluations" of Cultural Representation in AI
- 通过南亚工作坊,结合社区理解构建具身化评估框架
- 发现现有评估常脱离真实文化语境,导致误判与失真
- 适合关注文化公平与参与式AI评估的研究者与实践者
生成式AI在非西方文化表征方面的评估日益增多。然而,这些评估往往基于简化、抽象的表征理想,脱离了人们自身对文化表征的理解,忽视了文化表达固有的解释性与语境性。为此,我们提出“厚评估”(thick evaluations)概念,即一种更细致、情境化、对话式的评估框架,强调以社区自身的文化认知为基础。通过在南亚地区开展工作坊,研究人们如何解读和赋予自己文化相关的AI生成图像意义,我们发展出一套能体现社会世界表征复杂性的评估方法。该框架通过与社区共同构建评价指标,使测量标准与基层经验相契合,拓展了当前AI评估中关于表征的认知边界。
原文摘要 · Abstract (English)
Generative AI model outputs have been increasingly evaluated for their (in)ability to represent non-Western cultures. We argue that these evaluations often operate through reductive ideals of representation, abstracted from how people define their own representation and neglecting the inherently interpretive and contextual nature of cultural representation. In contrast to these 'thin' evaluations, we introduce the idea of 'thick evaluations:' a more granular, situated, and discursive measurement framework for evaluating representations of social worlds in AI outputs, steeped in communities' own understandings of representation. We develop this evaluation framework through workshops in South Asia, by studying the 'thick' ways in which people interpret and assign meaning to AI-generated images of their own cultures. We introduce practices for thicker evaluations of representation that expand the understanding of representation underpinning AI evaluations and by co-constructing metrics with communities, bringing measurement in line with the experiences of communities on the ground.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。