用言语行为理论重新定义生成模型的表征伤害,让抽象问题可分类可测量。
Taxonomizing Representational Harms using Speech Act Theory
- 基于言语行为理论,将语言模型行为分为不同施事类型,对应具体伤害。
- 提出新定义:刻板印象、贬低、抹除,形成细粒度伤害分类体系。
- 适合研究公平性、伦理评估与模型测评的研究者使用。
生成式语言系统造成的表征伤害被广泛认为是公平性问题的重要组成部分,但其定义通常不够明确。本文基于奥斯汀(Austin, 1962)的言语行为理论,提出一个理论框架,将生成系统引发的表征伤害概念化为特定施事行为(illocutionary acts)所产生的语用效果(perlocutionary effects,即现实影响)。在此基础上,结合语言人类学与社会语言学的相关研究,我们对刻板印象、贬低与抹除提出新的定义。进一步地,我们构建了一个细致的施事行为分类体系,超越了以往研究中的高层次分类。该框架还可支持开发有效的衡量工具。最后,通过案例研究,验证了本框架在回应近期关于表征伤害界定与测量的学术争论中的实用性。
原文摘要 · Abstract (English)
Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a theoretical contribution to the specification of representational harms by introducing a framework, grounded in speech act theory (Austin, 1962), that conceptualizes representational harms caused by generative language systems as the perlocutionary effects (i.e., real-world impacts) of particular types of illocutionary acts (i.e., system behaviors). Building on this argument and drawing on relevant literature from linguistic anthropology and sociolinguistics, we provide new definitions of stereotyping, demeaning, and erasure. We then use our framework to develop a granular taxonomy of illocutionary acts that cause representational harms, going beyond the high-level taxonomies presented in previous work. We also discuss the ways that our framework and taxonomy can support the development of valid measurement instruments. Finally, we demonstrate the utility of our framework and taxonomy via a case study that engages with recent conceptual debates about what constitutes a representational harm and how such harms should be measured.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。