arXiv:2504.15903cs.AI2025-04被引 2

测试噪声对大模型抽象推理能力的影响,发现当前模型极脆弱。

Impact of Noise on LLM-Models Performance in Abstraction and Reasoning Corpus (ARC) Tasks with Model Temperature Considerations

  • 在不同噪声和温度下系统评估多个大模型的推理表现
  • 噪声显著降低所有模型性能,即使强模型如GPT-4o也受影响
  • 适合关注模型鲁棒性与真实场景应用的研究者

大型语言模型(LLMs)在结构化推理任务中的表现引发广泛关注,尤其在抽象与模式识别方面。抽象与推理语料库(ARC)基准测试通过评估模型对新问题的泛化能力来衡量其推理水平。尽管GPT-4o在无噪声条件下能解决所有ARC任务,但DeepSeek R1和LLaMA 3.2却完全无法解决任何任务,表明其推理能力仍局限于简单模式匹配。为探究这一差距,我们系统地在不同噪声水平与温度设置下评估这些模型。结果表明,噪声的引入会持续削弱所有模型的表现,无论架构如何。这种下降凸显了当前大模型的共性缺陷:尽管展现出抽象推理迹象,但对输入扰动极为敏感。这种脆弱性引发了其在现实世界中应用的担忧,因噪声与不确定性普遍存在。通过对比不同架构对挑战的响应,本文揭示了现代大模型在推理任务中的结构性弱点。研究强调需发展更鲁棒、适应性强的AI系统,以应对现实场景中的模糊性与变异性,为未来提升模型泛化能力、鲁棒性与类人认知灵活性提供指导。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have generated growing interest in their structured reasoning capabilities, particularly in tasks involving abstraction and pattern recognition. The Abstraction and Reasoning Corpus (ARC) benchmark plays a crucial role in evaluating these capabilities by testing how well AI models generalize to novel problems. While GPT-4o demonstrates strong performance by solving all ARC tasks under zero-noise conditions, other models like DeepSeek R1 and LLaMA 3.2 fail to solve any, suggesting limitations in their ability to reason beyond simple pattern matching. To explore this gap, we systematically evaluate these models across different noise levels and temperature settings. Our results reveal that the introduction of noise consistently impairs model performance, regardless of architecture. This decline highlights a shared vulnerability: current LLMs, despite showing signs of abstract reasoning, remain highly sensitive to input perturbations. Such fragility raises concerns about their real-world applicability, where noise and uncertainty are common. By comparing how different model architectures respond to these challenges, we offer insights into the structural weaknesses of modern LLMs in reasoning tasks. This work underscores the need for developing more robust and adaptable AI systems capable of handling the ambiguity and variability inherent in real-world scenarios. Our findings aim to guide future research toward enhancing model generalization, robustness, and alignment with human-like cognitive flexibility.

大模型推理能力鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。