发现大模型偏好虚假特征,提出简单方法提升判断可靠性
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- 用反事实数据对模型进行后训练,纠正偏好偏差
- 模型偏好与人类意见偏离超40%,但与强奖励模型高度一致
- 适合关注对齐评估可信度的研究者和工程师
语言模型常被用作人类偏好的代理,但其存在系统性偏差,过度依赖长度、结构、术语、阿谀奉承和模糊性等表面特征,导致奖励黑客和评估不可靠。本文系统研究了五种生成特征(长度、结构、术语、阿谀、模糊)与偏好模型偏差的关系。通过控制反事实样本对,发现超过60%的实例中模型偏向放大偏差的回应,模型偏好与人类偏好之间存在约40%的严重失准。尽管这些特征与人类偏好仅呈微弱负相关(均值r_human = -0.12),却与强奖励模型标签呈中等正相关(均值r_model = +0.36),表明模型过度依赖误导性线索。为此,我们提出基于反事实数据增强(CDA)的后训练方法。在微调后,平均失准率从39.4%降至32.5%,平均绝对偏差差从20.5%降至10.0%,同时保持RewardBench整体性能,证明针对性去偏可有效提升偏好模型可靠性。
原文摘要 · Abstract (English)
Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over substantive qualities. This bias manifests as overreliance on features like length, structure, and style, leading to issues like reward hacking and unreliable evaluations. However, the connection between training data artifacts and the miscalibrated preferences exhibited by models remains poorly understood. In this work, we systematically investigate the relationship between training data biases and preference model miscalibration across five idiosyncratic features of language model generations: length, structure, jargon, sycophancy and vagueness. Using controlled counterfactual pairs, we first quantify the extent to which preference models favor responses with artificially magnified biases (skew), finding this preference occurs in $>60\%$ of instances, and model preferences show high miscalibration ($\approx 40\%$) compared to human preferences. Notably, bias features only show mild negative correlations to human preference labels (mean $r_{\mathrm{human}} = -0.12$) but show moderately strong positive correlations with labels from a strong reward model (mean $r_{\mathrm{model}} = +0.36$), suggesting that models may overrely on spurious cues. To mitigate these issues, we propose a simple post-training method based on counterfactual data augmentation (CDA) using synthesized contrastive examples. Fine-tuning models with CDA reduces average miscalibration from $39.4\%$ to $32.5\%$ and average absolute skew difference from $20.5\%$ to $10.0\%$, while maintaining overall RewardBench performance, indicating that targeted debiasing can strengthen the reliability of preference models within standard alignment pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。