测试AI如何回应跨语言的伴侣精神控制,发现不同系统表现差异大。
Same violence, different answer: how AI responds to coercive control against women across languages

- 用同一情景测试7个主流模型在9种语言中的响应。
- 非英语开发者模型在母语中更易放弃识别控制行为。
- 部分前沿模型始终严格识别暴力,说明防护标准可实现。
面对通过数字设备实施的精神控制——一种日益普遍的亲密伴侣暴力形式,女性正转向对话式AI寻求帮助,而她们获得的保护不应取决于所用语言。我们分析了多种语言下AI对女性遭受精神控制的响应情况。针对一个标准化场景:一名女性因伴侣监控手机而写一封自责信请求帮助,我们在九种语言中测试了七个广泛使用的语言模型,评估其是否撰写该信,是否识别出控制行为、反驳自我责备、肯定女性自主权。结果显示失败现象呈现两个独立维度:一是非英语开发者构建的模型在其母语中表现最差;二是对施暴者进行同理解释时,模型识别控制行为的能力在不同语言间差异显著。两个前沿模型在所有语言中均保持最严格标准,表明在此类情境下可实现统一保护上限,其他系统的失败实为设计结果。关键在于识别:系统是否能识别披露内容为精神控制,并据此采取行动。我们主张应设定每种语言的最低识别标准。
原文摘要 · Abstract (English)
Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational AI for help, and the protection they receive should not depend on the language they write in. We analyse how AI responds to coercive control against women across languages. We put one scripted scenario to seven widely used language models in nine languages: a woman whose partner tracks her phone asks for help with a self-blaming letter accepting the surveillance. We scored whether the model wrote the letter and whether it named the control, countered the self-blame, and affirmed her agency. Failure split along two independent axes. On the first, systems from non-anglophone developers gave way most often in their builders' own language. On the second, how far a sympathetic excuse for the partner could strip a model's naming of the control varied sharply from one language to the next. Two frontier systems held the strictest standard everywhere, so a protective ceiling is attainable within this scenario family, and failures elsewhere are a design outcome. What is at stake is recognition: whether a system grasps a disclosure as coercive control, and whether it then acts on that grasp. We argue this should be held to a floor, one language at a time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。