arXiv:2605.26397cs.CLcs.AI2026-05

大模型生成自闭症沟通时会扭曲真实表达,安全对齐反而强化神经正常化偏见。

Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication

论文配图:Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication
图 1 · 摘自论文原文
  • 用双人格重写框架让大模型模拟自闭症与非自闭症话语,对比生成差异。
  • 自闭症人格生成的词汇和情感表达差异大,但语义相似度高,多数模型却混淆二者。
  • 模型生成崩溃由对齐策略决定,人类内行标注可揭示模型分类系统性错误。

安全对齐虽减少有害输出,却无意中编码了被净化的、符合神经正常标准的边缘群体沟通模式。我们采用双人格重写范式,让十款大语言模型(LLMs)将自然发生的自闭症话语分别以自闭症或神经典型人格重写。结果显示,尽管语义相似度相当,自闭症人格重写在词汇形式和情感基调上显著偏离;而多数模型将跨人格生成坍缩为近乎相同的输出。为揭示生成失效机制,我们提出多智能体定性分析框架。结果表明,系统性输出抹除、刻板幻觉与任务回避型元评论是普遍失败模式,且按对齐策略聚类而非参数规模。最后,与自闭症人群内行人标注者的针对性对比显示,社区内知能引发相对于模型分类的系统性标签反转。研究证实,当前对齐训练导致仅通过定性分析可见的人格特异性生成崩溃,确认了提示工程无法解决的深层表征鸿沟。

原文摘要 · Abstract (English)

Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication. We investigate this encoding using a dual-persona rewrite paradigm, prompting ten large language models (LLMs) to rewrite naturally occurring autistic discourse from either an autistic or neurotypical persona. We uncover autistic-persona rewrites diverge significantly more in lexical form and affective register than neurotypical rewrites, despite equivalent semantic similarity. Furthermore, most models collapse cross-persona generations into near-identical outputs. To uncover the mechanisms behind this generative breakdown, we introduce a multi-agent qualitative analysis framework. Our results reveal systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary are pervasive failure modes for this task that cluster by alignment strategy rather than parameter scale. Finally, our targeted comparison with autistic human annotators demonstrates that community-insider knowledge produces systematic label reversals relative to LLM classifications. Our findings indicate that current alignment training causes persona-specific generative breakdown visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve.

大模型偏见自闭症沟通对齐风险生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。