arXiv:2410.14166cs.CLcs.AI2024-10

大模型在简单计数任务上表现差,但通过引导推理可显著提升准确率。

LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems

  • 设计多组实验验证主流失败假设,发现并非模型固有缺陷。
  • 引入推理引导后,模型在单词计数任务上准确率明显提升。
  • 适合关注模型能力迁移与推理机制优化的研究者参考。

大型语言模型在人类轻易完成的简单计数任务(如统计单词"strawberry"中字母'r'的数量)上仍表现不佳。现有观点普遍认为这源于分词方式、模型架构或训练数据,且难以避免。本文通过精心设计的多组评估设置,检验了这些主流假说的有效性,并考察了具备高级数学与编程推理能力的专用大模型向简单计数任务的能力迁移情况。尽管专用模型同样存在计数问题,但研究发现前述假说不成立,且揭示了从模型中激发有益知识的可能性。相较于微调或上下文学习等常见方法,我们证明引导推理是提升模型对任务感知力与响应准确性的最稳健高效手段。本研究呼吁重视模型能力获取与评估,强调在预训练中培养"先推理再作答"的意识。

原文摘要 · Abstract (English)

Interestingly, LLMs yet struggle with some basic tasks that humans find trivial to handle, e.g., counting the number of character r's in the word "strawberry". There are several popular conjectures (e.g., tokenization, architecture and training data) regarding the reason for deficiency of LLMs in simple word-based counting problems, sharing the similar belief that such failure stems from model pretraining hence probably inevitable during deployment. In this paper, we carefully design multiple evaluation settings to investigate validity of prevalent conjectures. Meanwhile, we measure transferability of advanced mathematical and coding reasoning capabilities from specialized LLMs to simple counting tasks. Although specialized LLMs suffer from counting problems as well, we find conjectures about inherent deficiency of LLMs invalid and further seek opportunities to elicit knowledge and capabilities from LLMs that are beneficial to counting tasks. Compared with strategies such as finetuning and in-context learning that are commonly adopted to enhance performance on new or challenging tasks, we show that engaging reasoning is the most robust and efficient way to help LLMs better perceive tasks with more accurate responses. We hope our conjecture validation design could provide insights into the study of future critical failure modes of LLMs. Based on challenges in transferring advanced capabilities to much simpler tasks, we call for more attention to model capability acquisition and evaluation. We also highlight the importance of cultivating consciousness of "reasoning before responding" during model pretraining.

大模型推理能力计数任务能力迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。