早期生成的词元置信度能有效预测多智能体推理质量
Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

- 用生成初期的词元概率作为推理质量的预测信号
- 前几词元的置信度比全程统计更准确预测推理质量
- 适合关注大模型推理可靠性的研究者参考
评估多智能体大模型系统中的推理质量极具挑战,尤其在无标准答案的开放任务中。本文探究解码过程中的词元级对数概率(即内在置信度)是否可作为推理质量的预测指标。基于辩论式作文评分框架,我们在两个ASAP作文数据集上比较了置信度代理指标与基于评分量规的裁判评分。结果表明,生成初期的词元置信度(尤其是前几词元)始终是推理质量最强的预测因子,优于全序列统计特征。对对数概率轨迹的分析显示,生成起始阶段最具异质性,信息量最大。此外,不同角色间存在系统性不对称:支持性推理中置信度与质量的匹配度高于对抗性批判。这些结果表明,早期解码动态为评估多智能体大模型推理可靠性提供了轻量化且有效的信号。
原文摘要 · Abstract (English)
Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from decoding, can predict reasoning quality as assessed by LLM-as-judge evaluation. Using a debate-based essay scoring framework, we compare confidence proxies against rubric-based judge scores across two ASAP essay sets. We find that early-token confidence, particularly within the first few generated tokens, is consistently the strongest predictor of reasoning quality, outperforming full-sequence statistics. Analysis of log-probability trajectories shows that the opening phase of generation is the most heterogeneous and therefore most informative. We also observe a systematic asymmetry between agent roles, with stronger alignment between confidence and quality for supportive reasoning than for adversarial critique. These results suggest that early decoding dynamics provide a lightweight and effective signal for estimating reasoning reliability in multi-agent LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。