用简单替换让大模型难辨自己输出,减轻自偏好问题。
Mitigating Self-Preference by Authorship Obfuscation
- 对候选答案做少量同义词替换,干扰模型识别自身输出
- 简单替换可显著降低自偏好,但彻底消除风格差异后偏见重现
- 揭示自识别可能存在于多层语义,完全解决仍具挑战
语言模型(LM)判官被广泛用于评估生成内容的质量。尽管优势明显,但其存在严重偏差,可能破坏评估的公正性。其中一种偏差是自偏好:模型倾向于选择自己的输出,而非其他模型或人类生成的内容。这种偏差难以消除,因为前沿的判官模型即使在未标注来源的情况下,也能区分自身输出。本文研究通过降低模型识别自身输出的能力来缓解自偏好。我们对成对比较中的候选答案施加黑盒扰动,通过作者身份混淆减少自识别。发现仅对少数词汇进行同义词替换即可有效降低自偏好。然而,当我们进一步扩大扰动以彻底中和候选答案间的风格差异时,自偏好现象又重新出现。结果表明,自识别与自偏好可在多个语义层面发生,即便初步效果良好,完全缓解仍面临根本性挑战。
原文摘要 · Abstract (English)
Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in evaluations. One such bias is self-preference: LM judges preferring their own answers over those produced by other LMs or humans. The bias is hard to eliminate as frontier LM judges can distinguish their own outputs from those of others, even when the evaluation candidates are not labeled with their sources. In this paper, we investigate strategies to mitigate self-preference by reducing the LM judges' ability to recognize their own outputs. We apply black-box perturbations to evaluation candidates in pairwise comparison to obfuscate the authorship and reduce self-recognition. We find that perturbations as simple as synonym replacement for a few words predictably reduce self-preference. However, we also uncover fundamental challenges to eliminating the bias: when we extrapolate our perturbations to a more complete neutralization of stylistic differences between the evaluation candidates, self-preference recovers. Our findings suggest that self-recognition and self-preference can happen on many semantic levels, and complete mitigation remains challenging despite promising initial results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。