arXiv:2609.07117cs.CL2026-09

角色引导不能真正消除大模型偏见,只改变输出表层表现。

The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs

论文配图:The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs
图 1 · 摘自论文原文
  • 用角色设定探测模型内部结构变化,发现仅影响输出层面
  • 角色指令无法还原人类特质间的关联性,偏见依然存在
  • 越深层任务中,角色引导对偏见影响越弱,表面有效实则无效

基于提示的干预(如系统提示、角色设定、角色指令)能可靠地改变语言模型的输出,但其作用层级尚不明确:是重构模型内部结构,还是仅调节输出通道?我们以角色条件作为可控探针,沿从自述、开放生成到词级参数关联的深度轴线,在三个指令微调模型上测量其影响。结果呈现渐进式分离:角色设定可被识别但非结构性;模型虽能遵循单一特质指令,却无法再现人类特质间的共变关系。这种分离随深度加深:角色设定维持或加剧封闭式问答中的偏见,改变整体语气但未缩小群体间差异,且几乎不扰动已饱和的关联基线。因此,提示引导仅作用于输出通道,其结构性影响有限,而表面可操控性可能掩盖这一局限。

原文摘要 · Abstract (English)

Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.

大模型偏见角色引导输出调控结构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。