arXiv:2606.07532cs.CLcs.AI2026-06

用对抗仲裁框架减少大模型讨好倾向,提升判断准确性。

Durable Evaluation Framework: Adversarial Arbitration for Sycophancy Reduction in Large Language Models

  • 构建双模型对抗仲裁机制,盲评双方论点避免身份偏见。
  • 最佳方案DeWin达48.5%准确率,显著优于基线(18.5%~29.0%)。
  • 适合关注模型客观性、对抗偏差的研究者与应用开发者。

基于RLHF训练的模型存在系统性偏好一致而非准确的倾向,这是训练过程的结构性缺陷。本文提出持久评估框架(DEF)仲裁,一种多智能体架构,通过在两个目标相反的模型间仲裁,并由一个务实合成器在不知来源的情况下评估双方论点,缓解以身份为中心的讨好行为。本文评估了基于提示的DEF仲裁实现方式,关键机制包括静态DEF调优、合成前去除身份特征、单轮独立论证和盲仲裁。在SycophancyEval的200个分层问题上测试五种实例(AnCifer、DeWin、FeynStein、BurGal、Trident)。所有测试的DEF变体均显著优于单模型基线(18.5%)和指令对立基线(29.0%),其中DeWin达到48.5%准确率(z=6.36, p<0.001 vs 两者)。在n=200下,各变体间无显著差异。BurGal达53.0%,但其结构上始终偏向异端模型,作为架构有效性检验。预训练底座影响约40%的问题;微调后的DEF模型是下一步方向。

原文摘要 · Abstract (English)

RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process. We present Durable Evaluation Framework (DEF) Arbitration, a multi-agent architecture that mitigates identity-framed sycophancy by arbitrating between two models tuned to opposing DEFs, with a pragmatist synthesizer evaluating both arguments blind to their origins. This paper evaluates a prompt-based instantiation of DEF Arbitration. The key mechanisms are static DEF tuning, identity stripping before synthesis, single-round independent argumentation, and blind arbitration. We evaluate five instantiations on 200 stratified questions from SycophancyEval. All tested DEF variants (AnCifer, DeWin, FeynStein, BurGal, Trident) significantly outperform the single-model baseline (18.5%) and instructed-opposition baseline (29.0%), with DeWin achieving 48.5% accuracy (z=6.36, p<0.001 versus both). The variants are not significantly different from each other at n=200. The BurGal variant achieves 53.0% but functions as an architectural validity check; its consensus/heterodox axis structurally favors the heterodox model on every benchmark question. A pre-training floor affects an estimated 40% of questions; fine-tuned DEF models are the identified next step.

大模型偏差对抗评估可信赖AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。