测试大模型解数独并解释思路,发现都难以给出有策略的推理过程。
Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku
- 用五种大模型解6×6数独并生成解释
- 无一模型能体现有策略的解题推理
- 适合研究AI可解释性与人机协作的学者
大型语言模型(LLMs)在人机协同决策中的成功依赖于其提供可信、渐进且定制化解释的能力。解复杂谜题如数独是此类协作的典型范例,其中清晰且个性化的解释往往比最终答案更重要。本研究评估了五种LLMs在求解和解释六六数独谜题方面的表现。尽管有一种模型在解题上表现有限,但没有任何模型能够以反映战略推理或直觉解题的方式解释求解过程。这些发现凸显了在大模型成为有效人机协同决策伙伴之前,仍需解决的重大挑战。
原文摘要 · Abstract (English)
The success of Large Language Models (LLMs) in human-AI collaborative decision-making hinges on their ability to provide trustworthy, gradual, and tailored explanations. Solving complex puzzles, such as Sudoku, offers a canonical example of this collaboration, where clear and customized explanations often hold greater importance than the final solution. In this study, we evaluate the performance of five LLMs in solving and explaining \sixsix{} Sudoku puzzles. While one LLM demonstrates limited success in solving puzzles, none can explain the solution process in a manner that reflects strategic reasoning or intuitive problem-solving. These findings underscore significant challenges that must be addressed before LLMs can become effective partners in human-AI collaborative decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。