用认知谜题测试大模型,区分机械记忆和真实推理能力。
Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles
- 设计双维度测评框架,分离故事熟悉度与推理复杂度。
- 模型对表面变化鲁棒,但在不对称场景中推理失败率超60%。
- 适合评估大模型真实推理能力,尤其关注认知状态追踪。
认知推理要求智能体基于部分观测和对其他智能体知识的了解推断世界状态。以往研究将大模型在认知谜题上的失败归因于记忆而非推理,我们认为这种二分法对新模型过于粗糙:记忆是模式化推理的一种极限情况,即模型将任务匹配到熟悉模板并应用对应解法。我们引入一个基于动态信念逻辑(DEL)的二维基准,分离叙事熟悉度与推理复杂度,从而区分模式化推理与认知推理。实验发现,模型对表面形式变化具有更强的鲁棒性,但当熟悉模式失效、需追踪碎片化认知状态时,表现持续不佳,表明其缺乏真正的认知推理能力。
原文摘要 · Abstract (English)
Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior work evaluating LLMs on epistemic puzzles often frames failures as memorization rather than reasoning. We argue that this dichotomy is too coarse for newer models: memorization is a limiting case of pattern-based reasoning, where a model matches a task to a familiar template and applies the corresponding solution. We introduce a two-dimensional benchmark over DEL-style puzzles, separating narrative familiarity from inference complexity, allowing us to distinguish pattern-based from epistemic reasoning. We find that models are substantially more robust to surface form changes than prior work suggested, yet consistently struggle in asymmetric settings where familiar patterns no longer apply and success requires tracking fragmented epistemic states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。