arXiv:2512.20812cs.CL2025-12

测试大模型在陌生符号下的计算能力,发现它们易受语义误导而出错。

Semantic Deception: When Reasoning Models Can't Compute an Addition

  • 用新符号替代数字和运算符,测试模型对抽象符号的处理能力
  • 4个大模型在简单计算中因语义干扰,准确率显著下降
  • 揭示模型依赖表面语义而非逻辑推理,适合关注AI决策可信度的研究者

大型语言模型(LLMs)在涉及人类价值观的决策任务中应用日益广泛。本文通过引入实验框架,测试模型在处理陌生符号时的推理能力。我们设计了语义欺骗:符号因形式或上下文产生误导性语义关联,以检验模型能否保持符号抽象性,还是依赖已学习的语义关联。将标准数字与数学运算符替换为新符号,要求模型完成以新记号表示的简单计算。实验针对4个LLM进行,结果表明,语义线索会显著降低模型在基础任务中的表现。这暴露出现有模型在符号操作上的局限性,表现出过度依赖表面语义的倾向,暗示思维链可能强化对统计相关性的依赖。即使模型看似遵循指令,语义干扰仍影响其基本能力。这些限制引发伦理与社会关切,挑战将推理能力归于模型的普遍趋势,警示在需要稳健符号推理的决策场景中,模型可能因训练中残留的语义关联而失败。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in situations where human values are at stake, such as decision-making tasks that involve reasoning when performed by humans. We investigate the so-called reasoning capabilities of LLMs over novel symbolic representations by introducing an experimental framework that tests their ability to process and manipulate unfamiliar symbols. We introduce semantic deceptions: situations in which symbols carry misleading semantic associations due to their form, such as being embedded in specific contexts, designed to probe whether LLMs can maintain symbolic abstraction or whether they default to exploiting learned semantic associations. We redefine standard digits and mathematical operators using novel symbols, and task LLMs with solving simple calculations expressed in this altered notation. The objective is: (1) to assess LLMs' capacity for abstraction and manipulation of arbitrary symbol systems; (2) to evaluate their ability to resist misleading semantic cues that conflict with the task's symbolic logic. Through experiments with four LLMs we show that semantic cues can significantly deteriorate reasoning models' performance on very simple tasks. They reveal limitations in current LLMs' ability for symbolic manipulations and highlight a tendency to over-rely on surface-level semantics, suggesting that chain-of-thoughts may amplify reliance on statistical correlations. Even in situations where LLMs seem to correctly follow instructions, semantic cues still impact basic capabilities. These limitations raise ethical and societal concerns, undermining the widespread and pernicious tendency to attribute reasoning abilities to LLMs and suggesting how LLMs might fail, in particular in decision-making contexts where robust symbolic reasoning is essential and should not be compromised by residual semantic associations inherited from the model's training.

大模型推理符号推理语义欺骗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。