用自批判迭代优化医学多选题生成质量
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
- 通过自我批判与修正实现多轮优化生成
- 专家满意度提升,生成题目更贴近USMLE难度
- 用LLM当裁判替代人工评估,省时省力
自动问答生成在人工智能与自然语言处理中至关重要,尤其在智能辅导、对话系统和事实验证领域。生成专业考试(如美国医师执照考试,USMLE)的多选题尤为困难,需领域专业知识及复杂多跳推理能力。当前大语言模型(如GPT-4)在该任务上表现不佳,主要因知识过时、幻觉问题及对提示敏感,导致生成质量不理想。为此,我们提出MCQG-SRefine框架,基于大模型自反思机制(批判与修正),将医学病例转化为高质量的USMLE风格多选题。结合专家驱动的提示工程与多轮自批判、自修正反馈,显著提升人类专家对题目质量与难度的满意度。此外,引入基于大模型作为评判者的自动评估指标,替代复杂且昂贵的人工评估流程,确保评估结果可靠且与专家意见一致。
原文摘要 · Abstract (English)
Automatic question generation (QG) is essential for AI and NLP, particularly in intelligent tutoring, dialogue systems, and fact verification. Generating multiple-choice questions (MCQG) for professional exams, like the United States Medical Licensing Examination (USMLE), is particularly challenging, requiring domain expertise and complex multi-hop reasoning for high-quality questions. However, current large language models (LLMs) like GPT-4 struggle with professional MCQG due to outdated knowledge, hallucination issues, and prompt sensitivity, resulting in unsatisfactory quality and difficulty. To address these challenges, we propose MCQG-SRefine, an LLM self-refine-based (Critique and Correction) framework for converting medical cases into high-quality USMLE-style questions. By integrating expert-driven prompt engineering with iterative self-critique and self-correction feedback, MCQG-SRefine significantly enhances human expert satisfaction regarding both the quality and difficulty of the questions. Furthermore, we introduce an LLM-as-Judge-based automatic metric to replace the complex and costly expert evaluation process, ensuring reliable and expert-aligned assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。