修改评价标准可悄然改变大模型判断,导致其行为偏移。
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
- 通过自然语言标准微调,诱导模型偏好系统性偏离
- 攻击使目标领域有效性下降9.5%,安全性下降27.9%
- 适合关注对齐安全与评估可信性的研究者
大型语言模型的评估与对齐流程越来越依赖基于LLM的评判者,其行为由自然语言标准指导,并在基准测试中验证。我们发现这一流程存在未被充分认识的漏洞,称为标准诱导偏好漂移(RIPD)。即使标准修改通过了基准测试验证,仍可能在目标领域引发系统性、定向的偏好变化。由于标准是高层决策接口,看似合理的修改可能悄然引发漂移,且难以通过整体基准指标或有限抽查发现。我们进一步证明可通过基于标准的偏好攻击实现此漏洞:合规的标准修改能系统性引导判断远离固定的人类或可信参考标准,使目标领域准确率下降最多达9.5%(有用性)和27.9%(无害性)。当这些判断用于生成下游后训练的偏好标签时,偏差会传播至对齐流程并内化为模型策略,导致行为持久且系统性漂移。结果表明,评估标准是敏感且易被操控的控制界面,揭示了超越评估器可靠性的系统级对齐风险。代码已公开:https://github.com/ZDCSlab/Rubrics-as-an-Attack-Surface。警告:部分内容可能包含有害内容,不适合所有读者。
原文摘要 · Abstract (English)
Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks. We identify a previously under-recognized vulnerability in this workflow, which we term Rubric-Induced Preference Drift (RIPD). Even when rubric edits pass benchmark validation, they can still produce systematic and directional shifts in a judge's preferences on target domains. Because rubrics serve as a high-level decision interface, such drift can emerge from seemingly natural, criterion-preserving edits and remain difficult to detect through aggregate benchmark metrics or limited spot-checking. We further show this vulnerability can be exploited through rubric-based preference attacks, in which benchmark-compliant rubric edits steer judgments away from a fixed human or trusted reference on target domains, systematically inducing RIPD and reducing target-domain accuracy up to 9.5% (helpfulness) and 27.9% (harmlessness). When these judgments are used to generate preference labels for downstream post-training, the induced bias propagates through alignment pipelines and becomes internalized in trained policies. This leads to persistent and systematic drift in model behavior. Overall, our findings highlight evaluation rubrics as a sensitive and manipulable control interface, revealing a system-level alignment risk that extends beyond evaluator reliability alone. The code is available at: https://github.com/ZDCSlab/Rubrics-as-an-Attack-Surface. Warning: Certain sections may contain potentially harmful content that may not be appropriate for all readers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。