用逻辑一致性检测大模型预测,实现即时评估。
Consistency Checks for Language Model Forecasters
- 基于套利思想设计通用一致性指标,检验预测逻辑是否自洽。
- 实验证明一致性指标与未来真实表现(Brier分数)高度相关。
- 提供2028年才揭晓的长期评测基准,适合研究者持续追踪。
预测任务难以评估,因真实结果需等待未来才能知晓。近期研究表明大语言模型在预测任务上的表现正快速接近人类水平,这引出一个关键问题:如何实现即时的模型评估?本文基于一致性检查框架,通过衡量模型在逻辑相关问题上的预测一致性来评估其性能。提出一种新的通用一致性度量方法——套利法:例如,若模型同时预测民主党与共和党赢得2024年美国总统选举的概率均为60%,则存在可被套利利用的逻辑矛盾。我们构建了一个自动化评估系统,自动生成基础问题,构造一致性检查,获取模型预测并测量其一致性。随后建立了一个符合严格评分规则的标准预测基准,并证明该(即时)一致性度量与未来真实表现的Brier分数具有显著相关性。此外,我们发布一个将于2028年才揭晓的长期一致性评测基准,为预测模型提供持续评估工具。
原文摘要 · Abstract (English)
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future. Recent work showing LLM forecasters rapidly approaching human-level performance begs the question: how can we benchmark and evaluate these forecasters instantaneously? Following the consistency check framework, we measure the performance of forecasters in terms of the consistency of their predictions on different logically-related questions. We propose a new, general consistency metric based on arbitrage: for example, if a forecasting AI illogically predicts that both the Democratic and Republican parties have 60% probability of winning the 2024 US presidential election, an arbitrageur can trade against the forecaster's predictions and make a profit. We build an automated evaluation system that generates a set of base questions, instantiates consistency checks from these questions, elicits the predictions of the forecaster, and measures the consistency of the predictions. We then build a standard, proper-scoring-rule forecasting benchmark, and show that our (instantaneous) consistency metrics correlate with LLM forecasters' ground truth Brier scores (which are only known in the future). We also release a consistency benchmark that resolves in 2028, providing a long-term evaluation tool for forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。