arXiv:2506.00064cs.CLcs.AI2025-06ACL被引 2

提出新基准,评估大模型主动纠错能力。

Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling

  • 构建新基准Mis-prompt,含四类任务与错误分类体系。
  • 实验发现现有大模型主动纠错能力差,微调可提升性能。
  • 适合研究大模型鲁棒性与自动纠错的学者使用。

大型语言模型在错误处理方面已取得显著进展。当前方法多为被动式,依赖明确的错误处理指令;但在真实场景中,此类指令往往缺失。本文提出主动错误处理的新挑战:如何在无显式指令下进行错误处理。为此,本文构建了新基准Mis-prompt,包含四个评估任务、错误类别分类体系及新数据集。进一步分析显示,当前大模型在主动错误处理上表现不佳,而基于错误处理样本的监督微调(SFT)能有效提升其能力。数据集将公开发布。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated significant advancements in error handling. Current error-handling works are performed in a passive manner, with explicit error-handling instructions. However, in real-world scenarios, explicit error-handling instructions are usually unavailable. In this paper, our work identifies this challenge as how to conduct proactive error handling without explicit error handling instructions. To promote further research, this work introduces a new benchmark, termed Mis-prompt, consisting of four evaluation tasks, an error category taxonomy, and a new evaluation dataset. Furthermore, this work analyzes current LLMs' performance on the benchmark, and the experimental results reveal that current LLMs show poor performance on proactive error handling, and SFT on error handling instances improves LLMs' proactive error handling capabilities. The dataset will be publicly available.

大模型错误处理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。