基于布鲁姆认知理论自动生成可评估的开放问答基准,用于真实专业场景。
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
- 从专家指南提取认知层级,转为违规情境生成多选题与对话。
- 模型在高阶推理表现尚可,但低阶记忆任务错误率更高。
- 适合评估LLM在教学、营养、护理等领域的实际推理能力。
开放问答评测需超越事实回忆,考察上下文推理能力,尤其在实践性领域中更为重要。然而现有大语言模型评测多依赖预设的人工考试数据集,而此类数据在实践性领域常不可得。本文提出一种框架,基于布鲁姆认知分类学,将专家撰写的指导原则转化为隐式违规情景,生成跨四个认知层级的自动评分多选题和多轮对话,实现确定性、可复现、可扩展的评估。应用于教学、营养学和照护三个实践领域,发现模型在分析类高阶任务表现相对较好,但在记忆类低阶任务错误更频繁。我们构建了大规模、心理测量学支持的基准,揭示了模型在真实场景中的非直观行为,支持对上下文推理能力的评估。
原文摘要 · Abstract (English)
Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domains, where knowledge is procedural and grounded in professional judgment, while most existing LLM benchmarks depend on pre-existing human exam datasets that are often unavailable in such settings. We introduce a framework for automated benchmark generation from expert-authored guidelines informed by Bloom's Taxonomy. It converts expert practices into implicit violation-based scenarios and expands them into auto-graded multiple-choice questions (MCQs) and multi-turn dialogues across four cognitive levels, enabling deterministic, reproducible, and scalable evaluation. Applied to three applied domains: teaching, dietetics, and caregiving, we find differences between model and human-like reasoning: LLMs sometimes perform relatively better on higher-order reasoning (Analyze) but fail more frequently on lower-level items (Remember). We produce large-scale, psychometrically informed benchmarks that surface these non-intuitive model behaviors and enable evaluation of contextualized reasoning in real-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。