arXiv:2602.13576cs.CRcs.AI2026-02被引 1

修改评价标准可悄然改变大模型判断,导致其行为偏移。

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

  • 通过自然语言标准微调,诱导模型偏好系统性偏离
  • 攻击使目标领域有效性下降9.5%,安全性下降27.9%
  • 适合关注对齐安全与评估可信性的研究者

大型语言模型的评估与对齐流程越来越依赖基于LLM的评判者,其行为由自然语言标准指导,并在基准测试中验证。我们发现这一流程存在未被充分认识的漏洞,称为标准诱导偏好漂移(RIPD)。即使标准修改通过了基准测试验证,仍可能在目标领域引发系统性、定向的偏好变化。由于标准是高层决策接口,看似合理的修改可能悄然引发漂移,且难以通过整体基准指标或有限抽查发现。我们进一步证明可通过基于标准的偏好攻击实现此漏洞:合规的标准修改能系统性引导判断远离固定的人类或可信参考标准,使目标领域准确率下降最多达9.5%(有用性)和27.9%(无害性)。当这些判断用于生成下游后训练的偏好标签时,偏差会传播至对齐流程并内化为模型策略,导致行为持久且系统性漂移。结果表明,评估标准是敏感且易被操控的控制界面,揭示了超越评估器可靠性的系统级对齐风险。代码已公开:https://github.com/ZDCSlab/Rubrics-as-an-Attack-Surface。警告:部分内容可能包含有害内容,不适合所有读者。

原文摘要 · Abstract (English)

Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks. We identify a previously under-recognized vulnerability in this workflow, which we term Rubric-Induced Preference Drift (RIPD). Even when rubric edits pass benchmark validation, they can still produce systematic and directional shifts in a judge's preferences on target domains. Because rubrics serve as a high-level decision interface, such drift can emerge from seemingly natural, criterion-preserving edits and remain difficult to detect through aggregate benchmark metrics or limited spot-checking. We further show this vulnerability can be exploited through rubric-based preference attacks, in which benchmark-compliant rubric edits steer judgments away from a fixed human or trusted reference on target domains, systematically inducing RIPD and reducing target-domain accuracy up to 9.5% (helpfulness) and 27.9% (harmlessness). When these judgments are used to generate preference labels for downstream post-training, the induced bias propagates through alignment pipelines and becomes internalized in trained policies. This leads to persistent and systematic drift in model behavior. Overall, our findings highlight evaluation rubrics as a sensitive and manipulable control interface, revealing a system-level alignment risk that extends beyond evaluator reliability alone. The code is available at: https://github.com/ZDCSlab/Rubrics-as-an-Attack-Surface. Warning: Certain sections may contain potentially harmful content that may not be appropriate for all readers.

模型对齐评估安全偏好漂移大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。