arXiv:2507.15707cs.CLcs.AI2025-07ACL被引 2

不同提问方式影响大模型推理准确率,且推理过程与最终答案不总一致。

Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked?

  • 对比多选、判断、长短回答三种提问方式对模型表现的影响
  • 推理步骤准确率与最终答案正确率常不匹配,存在分离现象
  • 选项数量和用词选择显著影响模型表现,提示提示工程关键性

大语言模型(LLMs)在不同类型的题目上进行了评估,包括选择题、是非题以及简答/长答。本研究探讨了提问方式对推理任务中模型准确性的影响。我们考察了五种主流大模型在三类问题上的表现,涵盖定量与演绎推理任务。性能指标包括推理步骤的准确率及最终答案的选择准确率。主要发现:(1) 不同提问方式下模型表现存在显著差异;(2) 推理过程准确率与最终答案正确率并不必然相关;(3) 选项数量及词语选择会影响模型表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been evaluated using diverse question types, e.g., multiple-choice, true/false, and short/long answers. This study answers an unexplored question about the impact of different question types on LLM accuracy on reasoning tasks. We investigate the performance of five LLMs on three different types of questions using quantitative and deductive reasoning tasks. The performance metrics include accuracy in the reasoning steps and choosing the final answer. Key Findings: (1) Significant differences exist in LLM performance across different question types. (2) Reasoning accuracy does not necessarily correlate with the final selection accuracy. (3) The number of options and the choice of words, influence LLM performance.

大模型推理评估提示设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。