用AI代理协作生成与评估高考数学选择题,发现生成题质量高但仍有差距。
Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation
- 人类协调多个AI代理完成题目生成全流程
- 生成题平均质量高,但难度与认知深度仍不足
- 适合关注AI科研协作与教育测评的从业者
大型语言模型正快速改变科学研究,但其对研究活动的实际影响仍缺乏实证。本研究通过混合方法评估了一种由人类协调的AI驱动研究流程:利用多个基于LLM的代理完成数据提取、语料构建、题库生成与评估。以生成和评估高考数学多选题(MCQ)为测试场景,共收集1,071道题目,通过代理从PDF中提取题目,将开放教材转换为结构化内容,将每道题与教材知识点对齐,按指定难度和认知层级生成新题,并使用24项标准的质量框架评估原始题与生成题。整体质量较高,但逐项分析与等效性检验显示生成题在技能深度、认知参与度、难度校准和元数据对齐方面存在持续差距;而语法流畅性、选项清晰度、无重复等表面质量表现稳定。此外,研究者角色转向规范设定、流程编排、验证与治理,包括设计规则、构建验证闭环、处理工具故障与溯源审计。研究揭示了未来科研中‘AI研究运维’能力的重要性。
原文摘要 · Abstract (English)
Advances in large language models (LLMs) are rapidly transforming scientific work, yet empirical evidence on how these systems reshape research activities remains limited. We report a mixed-methods pilot evaluation of an AI-orchestrated research workflow in which a human researcher coordinated multiple LLM-based agents to perform data extraction, corpus construction, artifact generation, and artifact evaluation. Using the generation and assessment of multiple-choice questions (MCQs) as a testbed, we collected 1,071 SAT Math MCQs and employed LLM agents to extract questions from PDFs, retrieve and convert open textbooks into structured representations, align each MCQ with relevant textbook content, generate new MCQs under specified difficulty and cognitive levels, and evaluate both original and generated MCQs using a 24-criterion quality framework. Across all evaluations, average MCQ quality was high. However, criterion-level analysis and equivalence testing show that generated MCQs are not fully comparable to expert-vetted baseline questions. Strict similarity (24/24 criteria equivalent) was never achieved. Persistent gaps concentrated in skill\ depth, cognitive engagement, difficulty calibration, and metadata alignment, while surface-level qualities, such as {grammar fluency}, {clarity options}, {no duplicates}, were consistently strong. Beyond MCQ outcomes, the study documents a labor shift. The researcher's work moved from ``authoring items'' toward {specification, orchestration, verification}, and {governance}. Formalizing constraints, designing rubrics, building validation loops, recovering from tool failures, and auditing provenance constituted the primary activities. We discuss implications for the future of scientific work, including emerging ``AI research operations'' skills required for AI-empowered research pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。