arXiv:2503.18394cs.LGcs.CL2025-03被引 28

通过外部重述提升大模型解情境谜题能力

Solving Situation Puzzles with Large Language Model and External Reformulation

  • 在多次问答后或错误猜测时,对谜题进行外部重述
  • 相比直接使用大模型,胜率更高,提问和猜测次数更少
  • 适合需要多轮交互推理的复杂问题求解场景

近年来,大语言模型在算术和符号推理任务中表现出色。然而我们发现,当涉及多轮对话的推理任务(如情境谜题)时,现有大模型(如ChatGPT)表现不佳,常重复提出细节过细或相似的问题,或在多次问答后仍作出错误猜测。为此,本文提出一种新型外部重述方法:在经历若干轮问答或模型出现错误猜测后,对原始情境谜题进行重构。实验表明,该方法在胜率、提问与猜测次数等指标上均优于直接使用大模型,凸显了战略性问题重述在增强大模型复杂交互推理能力方面的潜力。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have shown an impressive ability to perform arithmetic and symbolic reasoning tasks. However, we found that LLMs (e.g., ChatGPT) cannot perform well on reasoning that requires multiple rounds of dialogue, especially when solving situation puzzles. Specifically, LLMs intend to ask very detailed questions focusing on a specific aspect or same/similar questions after several rounds of Q&As. To help LLMs get out of the above dilemma, we propose a novel external reformulation methodology, where the situation puzzle will be reformulated after several rounds of Q&A or when the LLMs raise an incorrect guess. Experiments show superior performance (e.g., win rate, number of question/guess attempts) of our method than directly using LLMs for solving situation puzzles, highlighting the potential of strategic problem reformulation to enhance the reasoning capabilities of LLMs in complex interactive scenarios.

大模型推理情境谜题交互式推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。