arXiv:2510.24427cs.CL2025-10被引 6

用虚拟世界分离模型推理与记忆,精准评测其真实思维能力。

SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models

  • 构建真实与虚拟双世界并行数据集,保持推理难度一致
  • 发现模型在有记忆优势时性能提升15-20%,证明知识依赖存在
  • 适合评估大模型推理能力、研究记忆与思考的分离机制

评估语言模型(LMs)的推理能力受限于其庞大的参数化世界知识,基准表现常反映事实记忆而非真实推理。现有方法(如时间过滤、改写、对抗替换)无法清晰分离两者。本文提出SynthWorlds框架,将任务推理复杂度与事实知识解耦。构建两个结构相同的平行语料库:真实映射世界中模型可利用参数化知识,合成映射世界中此类知识无效。在此基础上设计两种镜像任务(多跳问答与页面导航),确保两世界推理难度相当。在仅参数化(如闭卷问答)和知识增强(如检索增强)设置下实验显示,模型始终存在显著的知识优势差距,即因记忆获得的性能提升达15%-20%。知识获取与整合机制虽可缩小差距,但无法消除,揭示系统改进空间。SynthWorlds全自动且可扩展,为语言模型提供此前难以实现的受控评估环境,支持精确、可验证的推理与记忆对比。

原文摘要 · Abstract (English)

Evaluating the reasoning ability of language models (LMs) is complicated by their extensive parametric world knowledge, where benchmark performance often reflects factual recall rather than genuine reasoning. Existing datasets and approaches (e.g., temporal filtering, paraphrasing, adversarial substitution) cannot cleanly separate the two. We present SynthWorlds, a framework that disentangles task reasoning complexity from factual knowledge. In SynthWorlds, we construct parallel corpora representing two worlds with identical interconnected structure: a real-mapped world, where models may exploit parametric knowledge, and a synthetic-mapped world, where such knowledge is meaningless. On top of these corpora, we design two mirrored tasks as case studies: multi-hop question answering and page navigation, which maintain equal reasoning difficulty across worlds. Experiments in parametric-only (e.g., closed-book QA) and knowledge-augmented (e.g., retrieval-augmented) LM settings reveal a persistent knowledge advantage gap, defined as the performance boost models gain from memorized parametric world knowledge. Knowledge acquisition and integration mechanisms reduce but do not eliminate this gap, highlighting opportunities for system improvements. Fully automatic and scalable, SynthWorlds provides a controlled environment for evaluating LMs in ways that were previously challenging, enabling precise and testable comparisons of reasoning and memorization.

模型评估推理能力知识分离语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。