WOMAC机制通过专家间互评提升竞赛公平性与预测可靠性。
WOMAC: A Mechanism For Prediction Competitions
- 用同行预测的最优聚合结果替代真实标签评分
- 实测显示其对专家真实表现的预测更可靠
- 适合标签噪声大的实际预测竞赛场景
竞赛广泛用于识别判断性预测与机器学习中的顶尖表现者,标准设计基于累积得分排名,但该方法既非激励相容也统计效率低。主要问题在于标签噪声导致弱者可能侥幸胜出,且赢家通吃的机制诱使专家虚报以提高胜率,即使降低期望得分。现有激励相容方案依赖随机机制,引入更多噪声且缺乏确定性。为此,本文提出新确定性机制WOMAC(最准确群体智慧),不将专家评分与有噪的真实结果对比,而是与给定噪声结果下同行预测的最佳聚合结果对比。该机制在典型场景中更具统计效率。尽管复杂度较高难以直接分析激励,本文提供了清晰理论基础。同时提供向量化高效实现,并在真实预测数据集上实证表明,相比标准机制,WOMAC能更可靠地预测专家的样本外表现。适用于标签噪声显著的各类竞赛场景。
原文摘要 · Abstract (English)
Competitions are widely used to identify top performers in judgmental forecasting and machine learning, and the standard competition design ranks competitors based on their cumulative scores against a set of realized outcomes or held-out labels. However, this standard design is neither incentive-compatible nor very statistically efficient. The main culprit is noise in outcomes/labels that experts are scored against; it allows weaker competitors to often win by chance, and the winner-take-all nature incentivizes misreporting that improves win probability even if it decreases expected score. Attempts to achieve incentive-compatibility rely on randomized mechanisms that add even more noise in winner selection, but come at the cost of determinism and practical adoption. To tackle these issues, we introduce a novel deterministic mechanism: WOMAC (Wisdom of the Most Accurate Crowd). Instead of scoring experts against noisy outcomes, as is standard, WOMAC scores experts against the best ex-post aggregate of peer experts' predictions given the noisy outcomes. WOMAC is also more efficient than the standard competition design in typical settings. While the increased complexity of WOMAC makes it challenging to analyze incentives directly, we provide a clear theoretical foundation to justify the mechanism. We also provide an efficient vectorized implementation and demonstrate empirically on real-world forecasting datasets that WOMAC is a more reliable predictor of experts' out-of-sample performance relative to the standard mechanism. WOMAC is useful in any competition where there is substantial noise in the outcomes/labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。