arXiv:2509.03647cs.CLcs.AI2025-09被引 3

用轻量级向量修正大模型自偏见,提升评估公平性

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

  • 通过激活空间的对比添加与优化构造纠偏向量
  • 非正当自偏好偏差降低最高达97%,显著优于提示和优化基线
  • 适用于需公正评估的模型调优与路由场景

大型语言模型日益作为自动化评估工具,但存在‘自偏好偏差’:倾向于偏好自身输出而非其他模型。这种偏差损害评估流程的公平性与可靠性,尤其在偏好微调与模型路由任务中。本文研究是否可通过推理阶段的轻量级引导向量缓解此问题而无需重新训练。构建了一个细分自偏好为正当与不正当两类的标注数据集,并采用两种方法生成引导向量:对比激活添加(CAA)与基于优化的方法。结果表明,引导向量可将不正当自偏好偏差减少高达97%,显著优于提示法与直接偏好优化基线。然而,该方法在正当自偏好与无偏一致性样本上表现不稳定,表明自偏好涉及多方向或非线性结构。这凸显了引导向量作为大模型评判保障的潜力与局限,也推动更鲁棒干预策略的发展。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97\%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.

大模型评估自偏好纠偏激活向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。