arXiv:2502.18099cs.LG2025-02NeurIPS被引 8

用博弈论方法减少人工标注,仅用2000条数据就让大模型对齐效果逼近主流水平。

Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment

  • 将对齐问题建模为领导者-追随者博弈,对抗标注噪声。
  • 仅用2000条人工标注数据,3轮迭代后胜过GPT-4多个基准测试。
  • 适合资源有限但需高效对齐大模型的研究或应用团队。

大语言模型对齐通常依赖大量精心标注的数据,成本高且易受标注噪声影响。本文提出基于斯塔克尔伯格博弈的偏好优化(SGPO),将对齐视为策略(领导者)与最坏情况偏好分布(追随者)之间的两阶段博弈。该框架在ε-Wasserstein球内保证$/mathcal{O}(ε)$有界后悔,具备形式化鲁棒性。我们进一步实现为斯塔克尔伯格自标注偏好优化(SSAPO),仅需少量人工标注的“种子”偏好,通过迭代自标注生成新提示。每轮中,采用分布鲁棒加权机制处理合成标注,防止噪声或偏差影响训练。令人惊讶的是,仅使用2000条种子标注(约为标准人类标注的1/30),SSAPO在三个迭代内即在多个基准上取得优于GPT-4的胜率。结果表明,严谨的斯塔克尔伯格建模可实现高效数据利用的大模型对齐,大幅降低对昂贵人工标注的依赖。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human preferences typically demands vast amounts of meticulously curated data, which is both expensive and prone to labeling noise. We propose Stackelberg Game Preference Optimization (SGPO), a robust alignment framework that models alignment as a two-player Stackelberg game between a policy (leader) and a worst-case preference distribution (follower). The proposed SGPO guarantees $\mathcal{O}(ε)$-bounded regret within an $ε$-Wasserstein ball, offering formal robustness to (self-)annotation noise. We instantiate SGPO with Stackelberg Self-Annotated Preference Optimization (SSAPO), which uses minimal human-labeled "seed" preferences and iteratively self-annotates new prompts. In each iteration, SSAPO applies a distributionally robust reweighting of synthetic annotations, ensuring that noisy or biased self-labels do not derail training. Remarkably, using only 2K seed preferences -- about 1/30 of standard human labels -- SSAPO achieves strong win rates against GPT-4 across multiple benchmarks within three iterations. These results highlight that a principled Stackelberg formulation yields data-efficient alignment for LLMs, significantly reducing reliance on costly human annotations.

大模型对齐自标注博弈论数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。