arXiv:2410.13029cs.CLcs.LG2024-10被引 5

测试GPT在无解数学题上是否该放弃回答,发现模型常乱猜。

When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems

  • 用常规解题提示测试GPT对无解题的应对能力
  • 模型在无解题上错误率超70%,却仍自信回答
  • 适合关注AI可靠性与推理安全的研究者

大型语言模型(LLMs)被广泛用于解决复杂的数学应用题。然而,这些模型容易产生幻觉,在面对无解问题时可能生成不准确的结果,带来潜在风险。尽管GPT模型已广泛应用且备受信任,但对其如何有效避免回答无解数学题以及如何提升拒答能力的研究仍不充分。本文通过在无解数学应用题(UWMP)数据集上使用典型可解题提示,评估GPT模型在无解情境下的表现。我们引入综合评价指标,涵盖拒答率、正确率和置信度三个维度。实验结果揭示了当前GPT模型在无解问题上的显著缺陷,存在严重幻觉现象,凸显出改进模型处理不确定性与复杂推理能力的迫切需求。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly relied upon to solve complex mathematical word problems. However, being susceptible to hallucination, they may generate inaccurate results when presented with unanswerable questions, raising concerns about their potential harm. While GPT models are now widely used and trusted, the exploration of how they can effectively abstain from answering unanswerable math problems and the enhancement of their abstention capabilities has not been rigorously investigated. In this paper, we investigate whether GPTs can appropriately respond to unanswerable math word problems by applying prompts typically used in solvable mathematical scenarios. Our experiments utilize the Unanswerable Word Math Problem (UWMP) dataset, directly leveraging GPT model APIs. Evaluation metrics are introduced, which integrate three key factors: abstention, correctness and confidence. Our findings reveal critical gaps in GPT models and the hallucination it suffers from for unsolvable problems, highlighting the need for improved models capable of better managing uncertainty and complex reasoning in math word problem-solving contexts.

大模型数学推理拒答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。