arXiv:2506.02259cs.GTcs.AI2025-06NeurIPS被引 2

提出更强的诚实激励机制,让真实反馈在任何单调效用下都更优。

Stochastically Dominant Peer Prediction

  • 用随机占优设计新评分规则,确保说真话总比撒谎更好
  • 新机制在二元信号下理论保证诚实最优,实测敏感度最高
  • 适合需要高可信反馈的AI对齐与噪声标签学习场景

获取可靠的人类反馈对机器学习任务至关重要,如处理噪声标签和对齐人工智能系统与人类偏好。同行预测机制通过比较个体与其他人的回答来评分,无需真实答案即可激励诚实报告。传统机制假设个体效用是得分的线性函数,但现实中非线性支付或效用更常见。本文提出随机占优诚实性(SD-truthfulness):说真话的得分分布在所有单调效用下严格优于其他策略。我们发现现有机制无法自然满足该性质。简单将得分四舍五入为二元抽奖可实现该性质,但会降低敏感性。通过更精细的舍入策略,可更好保留敏感性。此外,我们提出新的强制一致(EA)机制,在二元信号设置下于弱假设下理论保证SD-诚实性,并在实验中表现出所有已知SD-诚实机制中的最高敏感度。

原文摘要 · Abstract (English)

Eliciting reliable human feedback is essential for many machine learning tasks, such as learning from noisy labels and aligning AI systems with human preferences. Peer prediction mechanisms incentivize truthful reporting without ground truth verification by scoring agents based on correlations with peers. Traditional mechanisms, which ensure that truth-telling maximizes the expected scores in equilibrium, can elicit honest information while assuming agents' utilities are linear functions of their scores. However, in practice, non-linear payment rules are usually preferred, or agents' utilities are inherently non-linear. We propose stochastically dominant truthfulness (SD-truthfulness) as a stronger guarantee: the score distribution of truth-telling stochastically dominates all other strategies, incentivizing truthful reporting for a wide range of monotone utility functions. Our first observation is that no existing peer prediction mechanism naturally satisfies this criterion without strong assumptions. A simple solution -- rounding scores into binary lotteries -- can enforce SD-truthfulness, but often degrades sensitivity, a key property related to fairness and statistical efficiency. We demonstrate how a more careful application of rounding can better preserve sensitivity. Furthermore, we introduce a new enforced agreement (EA) mechanism that is theoretically guaranteed to be SD-truthful in binary-signal settings under mild assumptions, and empirically achieves the highest sensitivity among all known SD-truthful mechanisms.

同行预测激励机制诚实性随机占优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。