arXiv:2501.08203cs.CL2025-01中稿 · ACL被引 12

测试大模型在数学题中对乱加标点噪声的鲁棒性

ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving

  • 通过添加标点制造无信息损失的噪声输入
  • 所有模型在更多噪声下准确率下降,最差降超40%
  • 适合关注模型实际应用鲁棒性的研究者

尽管大型语言模型(LLMs)在数学问题求解任务中表现出色,但其对噪声输入的鲁棒性尚未得到充分研究。我们提出 ArithmAttack,用于评估当模型遇到包含额外标点符号噪声的提示时的鲁棒性。该方法易于实现且不造成信息丢失,因为未增删任何词语。我们在 GSM8K 和 MultiArith 数据集上评估了包括 LLama3、Mistral、Mathstral 和 DeepSeek 在内的八种 LLMs。实验表明,所有被测模型均对这类噪声敏感,噪声越多,性能越差。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have shown impressive capabilities in math problem-solving tasks, their robustness to noisy inputs is not well-studied. We propose ArithmAttack to examine how robust the LLMs are when they encounter noisy prompts that contain extra noise in the form of punctuation marks. While being easy to implement, ArithmAttack does not cause any information loss since words are not added or deleted from the context. We evaluate the robustness of eight LLMs, including LLama3, Mistral, Mathstral, and DeepSeek on noisy GSM8K and MultiArith datasets. Our experiments suggest that all the studied models show vulnerability to such noise, with more noise leading to poorer performances.

大模型评测数学推理鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。