测试大模型对堕胎污名的多层理解,发现其认知混乱且存在偏见。
Can LLMs Understand What We Cannot Say? Measuring Multilevel Alignment Through Abortion Stigma Across Cognitive, Interpersonal, and Structural Levels
- 用三层次量表测试627个角色对堕胎污名的理解能力。
- 模型低估认知羞耻感,高估人际担忧,且对年轻/非白人者错误赋值更高污名。
- 揭示模型内部矛盾:既说人人孤立却认为孤立者更不愿保密,适合政策与安全研究者参考。
随着大语言模型(LLMs)越来越多地介入具有污名化的健康决策,其对复杂心理现象的理解能力仍缺乏充分评估。我们探究了大模型是否能一致地理解堕胎污名在认知、人际和结构性三个层面的表现。通过使用经验证的个人层面堕胎污名量表(ILAS),系统测试了5个主流大模型在627个人口学多样角色中的表现,考察其在认知(自我评判)、人际(被评判与孤立担忧)和结构(社区谴责与披露模式)层面的表征。结果显示,所有维度上模型均未能通过真实理解测试:它们低估认知污名,高估人际污名,引入人口学偏差——对年轻、教育程度较低及非白人角色赋予更高的污名评分,并将秘密性视为普遍现象,尽管36%的人类受访者报告其实为开放态度。最关键的是,模型表现出内在矛盾:既高估孤立性,又预测孤立者更少保密,暴露其跨层次表征不一致。这些结果表明,现有对齐方法仅确保语言得体,而非多层级理解的一致性。本研究为大模型在多维心理建构上的理解缺陷提供了实证证据,警示高风险场景下必须采用新范式:设计(多层一致性)、评估(持续审计)、治理与监管(强制审计、问责机制、部署限制)以及公众对AI识别人类未言之隐的能力的认知教育。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) increasingly mediate stigmatized health decisions, their capacity to understand complex psychological phenomena remains inadequately assessed. Can LLMs understand what we cannot say? We investigate whether LLMs coherently represent abortion stigma across cognitive, interpersonal, and structural levels. We systematically tested 627 demographically diverse personas across five leading LLMs using the validated Individual Level Abortion Stigma Scale (ILAS), examining representation at cognitive (self-judgment), interpersonal (worries about judgment and isolation), and structural (community condemnation and disclosure patterns) levels. Models fail tests of genuine understanding across all dimensions. They underestimate cognitive stigma while overestimating interpersonal stigma, introduce demographic biases assigning higher stigma to younger, less educated, and non-White personas, and treat secrecy as universal despite 36% of humans reporting openness. Most critically, models produce internal contradictions: they overestimate isolation yet predict isolated individuals are less secretive, revealing incoherent representations. These patterns show current alignment approaches ensure appropriate language but not coherent understanding across levels. This work provides empirical evidence that LLMs lack coherent understanding of psychological constructs operating across multiple dimensions. AI safety in high-stakes contexts demands new approaches to design (multilevel coherence), evaluation (continuous auditing), governance and regulation (mandatory audits, accountability, deployment restrictions), and AI literacy in domains where understanding what people cannot say determines whether support helps or harms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。