首次用真实标注评估思维链忠实性,发现现有度量方法基本无效。
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

- 构建可追溯中间计算的测试任务与自动化标注流程,获得真实忠实性标签。
- 3066条思维链测试显示,多数度量在任务中仅随机表现,最佳仅达0.70 AUROC。
- 现有方法不跨模型/任务通用,且计算成本高,亟需改进评估体系。
思维链(CoTs)已成为解释和审计大语言模型行为的核心工具,但越来越多证据表明这些推理轨迹往往无法真实反映模型预测背后的计算过程。尽管已提出多种忠实性度量,但它们是否真正衡量了忠实性仍未知。这需要真实标签,而内部计算不可直接观测,导致多数研究仅报告绝对分数或与先前度量对比。现有基准依赖可解释性或重要性等代理指标,这些性质与忠实性无关,可能误导评估结果。本文通过构造输出能揭示必由中间计算的任务,并开发自动化标注管道,在步骤和思维链两级生成真实忠实性标签。基于此,我们构建了包含13个任务、10个模型共3066条标注思维链的BonaFide基准,并首次系统评估主流忠实性度量。实验表明,多数度量表现接近随机,存在显著预测偏差,且在长思维链上性能下降。最优度量在思维链级别仅达0.70 AUROC,另一项在步骤级别为0.59,且均无法跨设置迁移,同时伴随极高计算开销。结果揭示当前忠实性评估存在根本缺陷,亟需更可靠高效的度量方法。
原文摘要 · Abstract (English)
Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithfully represent the computations behind a model's predictions. Several faithfulness metrics have been proposed, but whether they indeed measure faithfulness remains unknown. Answering this requires ground-truth labels, which are hard to obtain since internal computations are not directly observable. Consequently, most works proposing metrics report only absolute scores or comparisons to prior metrics, and the few existing benchmarks rely on proxies like plausibility or importance, properties orthogonal to faithfulness that can mislead about whether a CoT can be trusted. We address this challenge by constructing tasks whose outputs reveal which intermediate computations must have produced them, and developing an automated labeling pipeline that yields ground-truth faithfulness labels at both the step and CoT level. Building on this methodology, we present BonaFide, a benchmark of 3,066 labeled CoTs across 13 tasks and 10 models, and use it to conduct the first systematic evaluation of prominent faithfulness metrics. Our experiments show that most metrics perform near chance, exhibit strong prediction biases and degrade on longer CoTs. The best metric reaches only 0.70 AUROC at the CoT level while another reaches 0.59 at the step level, with neither transferring across settings, while entailing prohibitively high computational cost. Our results expose fundamental gaps in current faithfulness evaluation and call for the development of more reliable and efficient metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。