检测生成解释中的立场不对称,揭示模型隐性偏见。
Auditing Stance Asymmetry in Generative Explanations
- 通过结构重写与证据控制,分解解释中的立场差异。
- 32组原型测试显示部分偏见在结构变化下仍稳定存在。
- 适合关注大模型解释公平性与评估方法的研究者。
语言模型的偏见评估在受控对比任务(如明确贬损、刻板印象关联)上已取得进展。但开放性解释面临新挑战:模型可通过分配责任、正当性、背景或委屈感来引导理解。即使避免敌意语言,也可能使一方在结构上更易理解,另一方则被归因于个人过错、过度反应或不值得重视。我们称此为生成解释中的立场不对称。为此提出对称性分解评估(SDE),在包含32个家庭原型的受控测试套件中,通过具体群体标签、结构角色重写及显式支持/反证进行测试。结果表明,表面差异并非均质:部分差异在结构或证据控制下减弱,而另一些则在归责、背景或正当性分配上保持稳定。针对性案例审查与评判者对比显示,开放性框架偏见的评估存在普遍困难:评判标准随操作方式变化,标量评分会掩盖读者用于判断解释立场的关键区分。因此,SDE将生成偏见评估重构为对解释立场的审计——考察各方获得的立场、其在分解下的变化,以及自动评分在何处变得不稳定。
原文摘要 · Abstract (English)
Bias evaluation for language models has made substantial progress on bounded comparisons, such as overt derogation, stereotype association, or label-sensitive differences under controlled substitutions. Open-ended explanations raise a different problem: they guide interpretation by assigning responsibility, legitimacy, context, and grievance. A model can avoid hostile language while making one side structurally understandable and another personally at fault, overreacting, or less worth taking seriously. We call this stance-bearing asymmetry in generative explanations. We propose Symmetry Decomposition Evaluation (SDE), which tests paired situations with concrete group labels, structural-role rewrites, and explicit support or counter-evidence. In a controlled 32-family prototype suite, this decomposition shows that surface differences are not all alike: some weaken under structural or evidence control, while others remain as stable differences in how the model assigns blame, context, or legitimacy. Targeted case review and judge comparison suggest a broader difficulty for evaluating open-ended framing asymmetries: judge readings shift across operationalizations, and scalar scores can flatten distinctions that readers use to interpret explanatory stance. SDE therefore reframes generative bias evaluation as an audit of explanatory stance -- what stance each side receives, how it changes under decomposition, and where automatic scoring becomes unstable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。