arXiv:2502.16399cs.IRcs.AI2025-02被引 1

用多模型协作提升AI评分准确性和可解释性。

Ensemble ToT of LLMs and Its Application to Automatic Grading System for Supporting Self-Learning

  • 通过多模型协同分析与模拟辩论优化答案
  • 在多个数据集上实现更高评分准确率
  • 适合需要高质量反馈的自主学习场景

为支持自主学习,及时提供详细评分反馈至关重要。现有基于大语言模型的评分系统大多依赖单一模型,性能受限。为此,我们提出集成式思维链(Ensemble Tree-of-Thought, Ensemble ToT)框架,通过整合多个模型增强输出效果。该框架包含三步:(1) 分析LLM表现,(2) 生成候选答案,(3) 优化形成最终结果。基于此,我们构建了一个评分系统:先评估各模型的评分倾向,再生成多份结果,最后通过模拟辩论进行融合。实验表明,该方法能有效协调多个大模型,实现更准确且可解释的自动评分。

原文摘要 · Abstract (English)

Providing students with detailed and timely grading feedback is essential for self-learning. While existing LLM-based grading systems are promising, most of them rely on one single model, which limits their performance. To address this, we propose Ensemble Tree-of-Thought (ToT), a framework that enhances LLM outputs by integrating multiple models. Using this framework, we develop a grading system. Ensemble ToT follows three steps: (1) analyzing LLM performance, (2) generating candidate answers, and (3) refining them into a final result. Based on this, our grading system first evaluates the grading tendencies of LLMs, then generates multiple results, and finally integrates them via a simulated debate. Experimental results demonstrate our approach's ability to provide accurate and explainable grading by effectively coordinating multiple LLMs.

自动评分多模型集成大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。