arXiv:2602.08259stat.MLcs.LG2026-02被引 4

用统计方法修正AI反馈偏差,让大模型对齐更准更快

A Statistical Framework for Alignment with Biased AI Feedback

  • 引入残差校正与密度比重加权,修正AI反馈系统性偏差
  • 在多个任务上逼近纯人工标注的对齐效果,性能提升显著
  • 适合大规模模型对齐,尤其需依赖AI评估的场景

现代对齐流程越来越多地用大型语言模型(LLM-as-Judge)替代昂贵的人类偏好标签。然而,相比高质量人类反馈数据集,AI标签可能存在系统性偏差。本文提出一种通用框架,支持异质提示-响应分布和外部人类反馈源,开发了两种去偏对齐方法:去偏直接偏好优化(DDPO)通过残差校正与密度比重加权,在保持DPO计算效率的同时缓解系统偏差;去偏身份偏好优化(DIPO)不假设参数化奖励模型,直接估计人类偏好概率。理论证明:DDPO是大规模对齐的实用高效解,DIPO则达到半参数效率界,具有统计最优性。在情感生成、摘要和单轮对话任务上的实证表明,所提方法显著提升对齐效率,并恢复接近基于完全人工标注数据训练的基准性能。

原文摘要 · Abstract (English)

Modern alignment pipelines are increasingly replacing expensive human preference labels with evaluations from large language models (LLM-as-Judge). However, AI labels can be systematically biased compared to high-quality human feedback datasets. In this paper, we develop two debiased alignment methods within a general framework that accommodates heterogeneous prompt-response distributions and external human feedback sources. Debiased Direct Preference Optimization (DDPO) augments standard DPO with a residual-based correction and density-ratio reweighting to mitigate systematic bias, while retaining DPO's computational efficiency. Debiased Identity Preference Optimization (DIPO) directly estimates human preference probabilities without imposing a parametric reward model. We provide theoretical guarantees for both methods: DDPO offers a practical and computationally efficient solution for large-scale alignment, whereas DIPO serves as a robust, statistically optimal alternative that attains the semiparametric efficiency bound. Empirical studies on sentiment generation, summarization, and single-turn dialogue demonstrate that the proposed methods substantially improve alignment efficiency and recover performance close to that of an oracle trained on fully human-labeled data.

模型对齐偏好学习去偏方法大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。