arXiv:2505.13636cs.LGcs.AI2025-05NeurIPS被引 9

通过同伴互评机制让大模型说真话,无需标注也不用微调。

Incentivizing Truthful Language Models via Peer Elicitation Games

  • 用多个基模型充当评判者,互相评分激励生成真实内容。
  • 实验显示在多个评测集上事实准确率显著提升,无需人工标签。
  • 理论保证模型长期收敛到诚实策略,适合追求可信生成的场景。

大型语言模型虽具强大生成能力,但仍易出现不一致和幻觉。本文提出无监督的同伴互评游戏(Peer Elicitation Games, PEG),通过一个生成器与多个来自不同基模型的判别器构成的互评机制实现对齐。判别器在同伴评估框架中互动,效用由基于行列式的互信息得分计算,可证明该机制在无需真实标签的情况下激励诚实报告。理论分析表明,各智能体通过在线学习可实现次线性遗憾,其累积表现趋近于事后最优固定诚实策略。此外,证明了最后迭代收敛至诚实纳什均衡,确保智能体实际策略随时间趋于稳定且真实。在多个基准上的实证评估显示,事实准确率显著提升。结果表明,PEG是一种无需监督或微调即可激发大模型诚实行为的实用方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework for aligning LLMs through a peer elicitation mechanism involving a generator and multiple discriminators instantiated from distinct base models. Discriminators interact in a peer evaluation setting, where utilities are computed using a determinant-based mutual information score that provably incentivizes truthful reporting without requiring ground-truth labels. We establish theoretical guarantees showing that each agent, via online learning, achieves sublinear regret in the sense their cumulative performance approaches that of the best fixed truthful strategy in hindsight. Moreover, we prove last-iterate convergence to a truthful Nash equilibrium, ensuring that the actual policies used by agents converge to stable and truthful behavior over time. Empirical evaluations across multiple benchmarks demonstrate significant improvements in factual accuracy. These results position PEG as a practical approach for eliciting truthful behavior from LLMs without supervision or fine-tuning.

大模型对齐诚实生成无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。