用荒诞世界测试大模型逻辑推理能力,发现其易受现实模式干扰。
Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities

- 将真实世界元素符号化后随机改造,生成逻辑一致但荒诞的场景。
- 多模型测试显示,多数模型在荒诞场景下推理准确率下降超40%。
- 适合评估大模型是否真会思考,而非依赖现实经验套模板。
尽管大型语言模型在各类任务中表现强大,其思维能力常受质疑,因它们有时无法解决人类可系统解决的问题。然而,现有研究多聚焦于用复杂问题破坏模型推理,而对简单逻辑推理的鲁棒性探索不足。本文提出Absurd World框架,通过改写真实世界的符号、动作、事件序列,构建逻辑自洽但违背常识的荒诞场景,以测试模型的逻辑推理能力。该方法自动生成多种荒诞设定,评估多个模型在简单与高级提示下的表现。结果表明,该框架能有效检验模型是否脱离现实模式进行真正推理。使用者可借此全面测试模型在真实任务变化下的稳健性。
原文摘要 · Abstract (English)
While extremely powerful and versatile at various tasks, the thinking capabilities of large language models (LLMs) are often put under scrutiny as they sometimes fail to solve problems that humans can systematically solve. However, recent literature focuses on breaking LLM reasoning with increasingly complex problems, and whether an LLM is robust in simple logical reasoning remains underexplored. This paper proposes Absurd World, a benchmarking framework, to test LLMs against altered realism, where scenarios are logically coherent, and humans can easily solve the tasks. Absurd World breaks a real-world model into symbols, actions, sequences, and events, which are automatically altered to create absurd worlds where the logic to solve the tasks remains the same. It evaluates a large collection of models with simple and advanced prompting techniques, and proves that it is an effective tool to determine LLMs' ability to think logically, ignoring the patterns learned from the real world. One can use this framework to extensively test an LLM against a real-world problem to verify whether the LLM's reasoning capability is robust against variations of the task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。