arXiv:2505.08389cs.CL2025-05

用凯撒密码构建抗污染评估基准,暴露大模型真实能力短板

Towards Contamination Resistant Benchmarks

  • 用简单凯撒密码设计抗污染评估基准
  • 多数大模型在无污染条件下仍表现不佳
  • 适合关注模型真实能力评估的研究者

大规模语言模型(LLMs)的快速发展重塑了自然语言处理领域。准确评估这些模型对理解其潜力和应对安全问题至关重要。然而,当前评估面临诸多挑战,其中数据污染尤为突出,严重削弱评估可靠性。本文提出‘抗污染性’概念,并设计一种基于凯撒密码(如位移为1时,'ab'变为'bc')的基准测试。尽管方法简单,但该基准具备强抗污染能力。我们在多种设置下对主流大模型进行测试,发现当控制污染后,这些模型在此基准上表现普遍不佳。结果揭示了现有大模型存在的问题,引发对其真实能力的深入思考。本工作推动了抗污染评估基准的发展,有助于实现更严格的模型评估,并提供关于大模型真实能力与局限性的新见解。

原文摘要 · Abstract (English)

The rapid development of large language models (LLMs) has transformed the landscape of natural language processing. Evaluating LLMs properly is crucial for understanding their potential and addressing concerns such as safety. However, LLM evaluation is confronted by various factors, among which contamination stands out as a key issue that undermines the reliability of evaluations. In this work, we introduce the concept of contamination resistance to address this challenge. We propose a benchmark based on Caesar ciphers (e.g., "ab" to "bc" when the shift is 1), which, despite its simplicity, is an excellent example of a contamination resistant benchmark. We test this benchmark on widely used LLMs under various settings, and we find that these models struggle with this benchmark when contamination is controlled. Our findings reveal issues in current LLMs and raise important questions regarding their true capabilities. Our work contributes to the development of contamination resistant benchmarks, enabling more rigorous LLM evaluation and offering insights into the true capabilities and limitations of LLMs.

大模型评估抗污染凯撒密码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。