显式监督可有效控制大模型的组合偏见,而基于偏好学习的方法失效。
Compositional Bias Control in Large Language Models: Preference Learning Fails, Supervision Succeeds
- 用显式正样本监督替代偏好学习来控制组合偏见
- 监督微调达99.87%约束合规率,且保持语言多样性
- 偏好优化方法仅4.53%有效,因无法表达逻辑连词
大型语言模型在中性职业语境下仍会生成性别刻板印象语言,反映深层社会偏见。现有方法包括提示工程、约束解码、后处理和微调对齐,但其效果与学习动态尚不明确。本文对比六种偏见缓解技术:仅提示、生成后过滤、基于DFA的Ctrl-G解码、监督微调(SFT)、直接偏好优化(DPO)和迭代零空间投影(INLP)。评估任务要求为20个Winogender衍生职业生成包含至少一个主动性和共情性描述的句子。通过约束符合度、词汇多样性和流畅性量化控制强度与自然度的权衡。结果显示:SFT实现99.87%±0.15%合规率,且保持高词汇多样性;而DPO虽训练稳定,却仅达4.53%±0.82%;Ctrl-G保证完全合规,但严重损害流畅性与多样性。原因在于偏好信号编码排序而非逻辑合取,无法满足组合约束。唯有显式正向监督能有效缓解组合偏见,凸显偏好学习在逻辑结构泛化上的局限,强调显式监督对公平流畅可控生成的必要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) still produce gender-stereotyped language even in occupation-neutral contexts that reflect deep societal biases (Rudinger et al., 2018). To address this, prior work has proposed prompting, constrained decoding (Dathathri et al., 2020; Zhou et al., 2024), post-processing, and fine-tuning-based alignment (Rafailov et al., 2023; Ravfogel et al., 2022). However, the comparative efficacy and learning dynamics remain little understood. We report a comparative analysis of six control techniques for bias mitigation: prompt-only, generate-and-filter, DFA-based Ctrl-G decoding, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Iterative Nullspace Projection (INLP). We evaluate each method on a compositional constraint task. This task requires generating sentences that contain at least one agentic and one communal descriptor for each of the twenty Winogender-derived occupations. We quantify trade-offs between control strength and naturalness with evaluations of constraint compliance, lexical diversity, and fluency. Our results reveal key contrasts among the methods: SFT achieves 99.87 +- 0.15% compliance and high lexical diversity, while DPO, despite similar training stability, fails at 4.53 +- 0.82%. Ctrl-G guarantees perfect compliance, but at the cost of severely reduced fluency and diversity. Preference-based learning fundamentally differs: it cannot satisfy compositional constraints, as binary preference signals encode ranking, not logical conjunctions. Only explicit positive supervision enables mitigation of compositional biases; preference-based alignment fails to generalize logical structures, underscoring the limitations of preference learning and the necessity of explicit supervision for fair and fluent controlled generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。