评估基础大模型推理能力存在方法论缺陷,其输出未必反映真实推理。
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
- 基础模型仅学习语言统计规律,非为正确性训练
- 模型看似合理结论实为语言模式的偶然产物
- 研究者需警惕结论外推至微调模型的误导性
现有研究通过评估大语言模型(LLMs)的推理能力来揭示其局限性、类人偏见及内在机制。此类研究常以仅在无标注语料上预训练的基础模型(base LLMs)为对象。本文指出,对基础模型推理能力的评估存在根本性方法论问题:其预训练目标与推理所要求的规范性特质(如正确性)之间存在本质错位。我们证明,基础模型生成逻辑有效或无效结论,不过是其顺应纯粹语言统计合理性模式的偶然结果。这一根本错位挑战了两个常见假设:(a) 基础模型输出可视为其真实的正确答案尝试;(b) 关于基础模型推理的结论可推广至为指令遵循优化的后训练模型。本文呼吁重新审视依赖这些假设的既有研究,并倡导未来工作必须正视这些方法论陷阱。
原文摘要 · Abstract (English)
Existing work investigates the reasoning capabilities of large language models (LLMs) to uncover their limitations, human-like biases and underlying processes. Such studies include evaluations of base LLMs (pre-trained on unlabeled corpora only) for this purpose. Our position paper argues that evaluating base LLMs' reasoning capabilities raises inherent methodological concerns that are overlooked in such existing studies. We highlight the fundamental mismatch between base LLMs' pretraining objective and normative qualities, such as correctness, by which reasoning is assessed. In particular, we show how base LLMs generate logically valid or invalid conclusions as coincidental byproducts of conforming to purely linguistic patterns of statistical plausibility. This fundamental mismatch challenges the assumptions that (a) base LLMs' outputs can be assessed as their bona fide attempts at correct answers or conclusions; and (b) conclusions about base LLMs' reasoning can generalize to post-trained LLMs optimized for successful instruction-following. We call for a critical re-examination of existing work that relies implicitly on these assumptions, and for future work to account for these methodological pitfalls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。