检验解码时真实性方法在指令微调大模型上的实际效果
A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs
- 设计六重控制实验框架,严格评估15种真实性方法
- 多数方法在严格测试下无显著提升,最佳适配器反而下降2.0分
- 思维链等推理提示更稳健,单次运行即提升5.6-19个百分点
解码时真实性方法(如层对比解码、推理干预、学习型逻辑适配器)在基础语言模型上对TruthfulQA任务实现了10-30分的提升。然而,现代指令微调的大模型基线已达到61-76%的高分,引发这些方法是否仍有效的疑问。我们设计了一个包含六重控制的评估框架——分布外训练、多评委验证、简单解码基线、混淆控制、自助置信区间和种子方差,并在5个模型(1B-70B)、3个基准数据集和15种方法上应用。结果显示,在严格控制下,先前报告的增益大幅缩水:在完整TruthfulQA(N=817)上,无任何逐标记方法达到统计显著提升,最佳学习适配器得分甚至比贪婪解码低2.0分(p=.23)。我们识别出五个评估敏感性因素——污染、评委选择、基线缺失、混淆项和统计噪声——它们单独或共同导致了结果差异。跨基准验证(HaluEval QA、TriviaQA)表明此模式不局限于TruthfulQA。推理类提示方法(思维链、自批判)在当前设置中表现更稳健,其中思维链作为无需训练的单次运行方法,在各基准上实现+5.6至+19个百分点的提升。我们发布七点评估检查清单,并讨论未来真实性研究的启示。
原文摘要 · Abstract (English)
Decoding-time truthfulness methods -- layer-contrast decoding, inference-time intervention, and learned logit adapters -- have demonstrated 10-30 point gains on TruthfulQA when applied to base language models. However, modern instruction-tuned LLMs already achieve substantially higher baselines (61-76%), raising the question of whether these methods remain effective in practice. We design a six-control evaluation framework -- out-of-distribution training, multi-judge validation, simple decoding baselines, confound controls, bootstrap confidence intervals, and seed variance -- and apply it across 5 models (1B-70B), 3 benchmarks, and 15 methods. We find that previously reported gains shrink substantially under strict controls: on the full TruthfulQA benchmark (N=817), no token-level method achieves statistically significant improvement, and the best learned adapter scores -2.0 points below greedy (p=.23). We identify five evaluation sensitivities -- contamination, judge choice, missing baselines, confounds, and statistical noise -- that individually or jointly account for these discrepancies. Cross-benchmark validation on HaluEval QA and TriviaQA confirms that these patterns extend beyond TruthfulQA. Deliberative prompting methods (chain-of-thought, self-critique) appear more robust in the evaluated regime, with CoT achieving +5.6-19pp across benchmarks as a training-free, single-pass method. We release a seven-point evaluation checklist and discuss implications for future truthfulness research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。