答案提取方式影响大模型推理评估,新方法提升结果可靠性。
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
- 用额外推理重生成答案,减少提取规则干扰。
- 在数学和开放问答任务中性能更优,结果更稳定。
- 适合关注评估公平性与结果可信度的研究者。
评估生成模型(如大语言模型)通常采用问答任务,最终答案基于选项概率选择。然而,对于需要推理的模型,答案提取方式至关重要。我们的研究发现,推理模型的性能及其最终答案分布对所用答案提取算法极为敏感。为此,我们提出一种基础框架:答案重生成。该方法通过一次额外的模型推理,以提示'Answer:'开头,将原始输入与输出重新生成,再从中选取或提取最终答案。该提取规则无关的方法展现出更优性能与更强鲁棒性。我们已将其应用于通用数学问题与开放问答任务。分析与该框架可为模型评估提供更可靠的依据。
原文摘要 · Abstract (English)
Evaluating generative models, such as large language models (LLMs), commonly involves question-answering tasks where the final answer is selected based on probability of answer choices. On the other hand, for models requiring reasoning, the method of answer extraction plays a critical role. Our research reveals that the performance of reasoning models and their final answer distributions are highly sensitive to the answer extraction algorithm employed. In order to mitigate this, we propose a basic framework: Answer Regeneration. The method uses an additional model inference, providing the prior input and output prefaced by the prompt "Answer:". The final answer is then selected or extracted from the regenerated output. We show that this extraction-rule-agnostic approach exhibits improved performance and enhanced robustness. Furthermore, we have applied this framework to general math problems and open-ended question answering tasks. Our analysis and this framework could offer a more reliable results for model evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。