arXiv:2506.04535cs.CL2025-06

测试大模型对无解问题的回答能力,发现表现远未完美。

BSBench: will your LLM find the largest prime number?

  • 设计新基准测试无解问题的推理能力
  • 现有模型在无解问题上准确率远低于100%
  • 适合评估模型抗误导与逻辑边界能力

我们提出,对大语言模型进行无合理答案的问题测试,并非毫无意义。为此,我们构建了一个基准测试,可评估模型在无解问题上的表现,并提出方法修改现有数据集以适配此类测试。实验发现,当前主流模型在这些无解问题上的表现远未达到理想水平。相关代码与数据已开源于https://github.com/L3G5/impossible-bench。

原文摘要 · Abstract (English)

We propose that benchmarking LLMs on questions which have no reasonable answer actually isn't as silly as it sounds. We also present a benchmark that allows such testing and a method to modify the existing datasets, and discover that existing models demonstrate a performance far from the perfect on such questions. Our code and data artifacts are available at https://github.com/L3G5/impossible-bench

大模型评测无解问题基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。