arXiv:2505.11341cs.CL2025-05EMNLP被引 5

构建首个大规模批判性问题生成数据集,推动大模型批判性思维评估

Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models

  • 构建5000条人工标注的批判性问题数据集
  • 提出与人类判断高度相关的自动评估方法
  • 11个大模型零样本测试展现任务挑战性

批判性问题生成(CQs-Gen)旨在通过生成揭示隐含假设和质疑论证有效性的提问,促进系统批判性思维。尽管该领域关注度上升,但进展受限于缺乏合适数据集和自动评估标准。本文首次构建包含约5000条人工标注问题的大规模数据集,并研究自动评估方法,提出基于参考答案的评估策略在与人类判断相关性上表现最佳。对11个大语言模型的零样本评估建立了强基线,凸显任务难度。论文提供数据、代码及公开排行榜,以推动模型性能提升,并探索该任务在自动化推理与人类批判性思维培养中的实际价值。

原文摘要 · Abstract (English)

The task of Critical Questions Generation (CQs-Gen) aims to foster critical thinking by enabling systems to generate questions that expose underlying assumptions and challenge the validity of argumentative reasoning structures. Despite growing interest in this area, progress has been hindered by the lack of suitable datasets and automatic evaluation standards. This paper presents a comprehensive approach to support the development and benchmarking of systems for this task. We construct the first large-scale dataset including ~5K manually annotated questions. We also investigate automatic evaluation methods and propose reference-based techniques as the strategy that best correlates with human judgments. Our zero-shot evaluation of 11 LLMs establishes a strong baseline while showcasing the difficulty of the task. Data and code plus a public leaderboard are provided to encourage further research, not only in terms of model performance, but also to explore the practical benefits of CQs-Gen for both automated reasoning and human critical thinking.

批判性思维大模型评估数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。