arXiv:2504.05693cs.CLcs.AI2025-04

用多轮迭代LLM评估,自动提升题目质量判断准确性。

STRIVE: A Think & Improve Approach with Iterative Refinement for Enhancing Question Quality Estimation

  • 多LLM协同生成多份评估,选最优结果
  • 通过迭代评审使评估指标趋于稳定,相关性提升
  • 适合教育AI、智能出题系统研发者参考

自动评估题目质量对教育者至关重要,可节省时间、保证一致性并即时反馈以优化教学材料。我们提出一种新方法STRIVE(基于多LLM的结构化思考与迭代改进,用于提升验证性题目评估),利用一系列大型语言模型实现自动题目评价。该方法通过生成多个基于题目优缺点的评估,并选取其中最佳方案来估算题目质量;随后由另一LLM进行迭代审查与回应,直至评估指标收敛。此复杂流程显著提升了题目质量评估的准确性与深度,支持多样化学习者并促进教育实践。相关性分析显示,相比基线方法,本方法与人工判断的相关性更高;误差分析表明,如相关性与适宜性等指标在使用STRIVE后显著优于人工判断。

原文摘要 · Abstract (English)

Automatically assessing question quality is crucial for educators as it saves time, ensures consistency, and provides immediate feedback for refining teaching materials. We propose a novel methodology called STRIVE (Structured Thinking and Refinement with multiLLMs for Improving Verified Question Estimation) using a series of Large Language Models (LLMs) for automatic question evaluation. This approach aims to improve the accuracy and depth of question quality assessment, ultimately supporting diverse learners and enhancing educational practices. The method estimates question quality in an automated manner by generating multiple evaluations based on the strengths and weaknesses of the provided question and then choosing the best solution generated by the LLM. Then the process is improved by iterative review and response with another LLM until the evaluation metric values converge. This sophisticated method of evaluating question quality improves the estimation of question quality by automating the task of question quality evaluation. Correlation scores show that using this proposed method helps to improve correlation with human judgments compared to the baseline method. Error analysis shows that metrics like relevance and appropriateness improve significantly relative to human judgments by using STRIVE.

教育AILLM评估迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。