arXiv:2606.30987cs.CLecon.GN2026-06

用大模型分析预测解释,能准确判断预测质量。

Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments

论文配图:Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
图 1 · 摘自论文原文
  • 设计60种推理模式,由大模型自动评分解释质量。
  • 在5.5万组预测数据中,预测准确率优于传统方法。
  • 适合评估预测者水平,尤其擅长识别表现差的预测者。

决策者常依赖带文字解释的专家判断,但解释质量难以规模化评估。预测竞赛提供了天然测试场景:概率性预测搭配自然语言理由,并根据实际结果打分。我们提出解释质量标记(EQMs),一组基于理论的60种推理模式,由大语言模型进行评分。在为期多年、超过5.5万对预测-解释数据的预注册分析中,EQMs在预测和预测者层面均能有效预测准确性,显著优于传统的文本分析方法。超过90%的模式级EQM-准确率相关性与研究假设方向一致。信号具有不对称性:EQMs更可靠地识别可能表现不佳的预测者,而非区分顶尖预测者。与传统预测能力指标对比,EQMs在预测层面是最强预测因子,在预测者层面也具竞争力,但弱于历史准确率。人工对解释质量的评分与准确率的相关性较弱,且过度依赖解释长度。结果在独立预测研究中得到验证。EQMs提供了一种可扩展、可解释的方法,从书面解释中提取与判断相关的信息。

原文摘要 · Abstract (English)

Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a natural testing ground: probabilistic judgments are paired with natural-language rationales and scored against realized outcomes. We introduce Explanation Quality Markers (EQMs), a set of sixty theory-guided reasoning patterns scored by large language models (LLMs). In a pre-registered analysis of over 55,000 forecast-rationale pairs from a multiyear forecasting tournament, EQMs predict accuracy at both the forecast and forecaster levels, consistently outperforming pre-LLM text-analysis methods. More than 90% of statistically significant pattern-level EQM-accuracy correlations match our directional hypotheses. The signal is asymmetric: EQMs identify likely underperformers more reliably than they distinguish the very best forecasters. Benchmarked against traditional indicators of forecasting skill, EQMs are the strongest predictor at the forecast level and competitive at the forecaster level, though weaker than prior accuracy. Human ratings of rationale quality are less consistently correlated with accuracy and place disproportionate weight on rationale length. Results transfer to an independent forecasting study. EQMs provide a scalable, interpretable method for extracting judgment-relevant information from written explanations.

预测评估解释质量大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。