arXiv:2505.17760cs.LGcs.AI2025-05被引 3

用引导向量让大模型评委更识破欺骗性回答。

But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors

  • 通过单个示例训练诚实引导向量,生成对比样本
  • GPT-4.1和Claude Haiku的判别准确率提升至0.946和0.929
  • 适合用于挑战性任务中的白盒评估与诚信审计

大模型作为评判者被广泛用于替代人工评估,但现有方法依赖黑箱访问,难以识别微妙的不诚实行为(如阿谀奉承和操纵)。我们提出JUSSA框架,利用模型内部表征,从单一训练样本中优化出促进诚实性的引导向量,生成对比性替代响应,为评判者提供识别不诚实行为的参照。我们在一个新型操纵性基准上进行测试,该基准包含人类验证的多级不诚实响应对,结果显示,无论是GPT-4.1(AUROC从0.893提升至0.946)还是Claude Haiku(从0.859提升至0.929),判别性能均有显著提升。然而当任务复杂度超出评判者能力范围时性能下降,表明对比评估在任务具有挑战性但仍在评判者能力范围内时最有效。层分析显示,中间层的表征差异对引导最为敏感。本工作证明引导向量可作为评估工具而非仅用于输出优化,为全面白盒审计开辟新路径。

原文摘要 · Abstract (English)

LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycophancy and manipulation. We introduce Judge Using Safety-Steered Alternatives (JUSSA), a framework that leverages a model's internal representations to optimize an honesty-promoting steering vector from a single training example, generating contrastive alternatives that give judges a reference point for detecting dishonesty. We test JUSSA on a novel manipulation benchmark with human-validated response pairs at varying dishonesty levels, finding AUROC improvements across both GPT-4.1 (0.893 $\to$ 0.946) and Claude Haiku (0.859 $\to$ 0.929) judges, though performance degrades when task complexity is mismatched to judge capability, suggesting contrastive evaluation helps most when the task is challenging but within the judge's reach. Layer-wise analysis further shows that steering is most effective in middle layers, where model representations begin to diverge between honest and dishonest prompt processing. Our work demonstrates that steering vectors can serve as tools for evaluation rather than for improving model outputs at inference, opening a new direction for thorough white-box auditing.

大模型评估诚实检测引导向量白盒审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。