用大模型评估教育资料相关性,效果比传统方法好得多
Validating LLM-Generated Relevance Labels for Educational Resource Search
- 用教育领域特化提示词让大模型判断资源相关性
- 最佳方案与人工评判一致度达κ=0.65,显著优于基础方法
- 适合教育技术研究者和智能教学系统开发者
信息检索中的手动相关性判断成本高且需专业背景,促使研究转向使用大语言模型(LLMs)进行自动评估。尽管在通用网络搜索中已有进展,但针对教育类资源这类特定领域的有效性尚未明确。本研究通过教学专业人士的用户研究,收集并发布了包含401条人工相关性判断的数据集。比较了三种提示结构:基于以往工作的两维度基线、源自教育文献的12维评价体系,以及由研究参与者直接反馈制定的标准。采用领域特化框架后,大模型与人工判断的一致性显著提升(Cohen's κ最高达0.650),尤其参与者自定义的5维和10维框架表现稳健,GPT-3.5分别达到κ=0.613和0.639。系统级评估显示,大模型判断能可靠识别出最优检索方法(RBO 0.71–0.76),同时保持合理区分度(RBO 0.52–0.56)。结果表明,在恰当提示下,大模型可有效评估教育类资源相关性,但性能受框架复杂度与输入结构影响。
原文摘要 · Abstract (English)
Manual relevance judgements in Information Retrieval are costly and require expertise, driving interest in using Large Language Models (LLMs) for automatic assessment. While LLMs have shown promise in general web search scenarios, their effectiveness for evaluating domain-specific search results, such as educational resources, remains unexplored. To investigate different ways of including domain-specific criteria in LLM prompts for relevance judgement, we collected and released a dataset of 401 human relevance judgements from a user study involving teaching professionals performing search tasks related to lesson planning. We compared three approaches to structuring these prompts: a simple two-aspect evaluation baseline from prior work on using LLMs as relevance judges, a comprehensive 12-dimensional rubric derived from educational literature, and criteria directly informed by the study participants. Using domain-specific frameworks, LLMs achieved strong agreement with human judgements (Cohen's $κ$ up to 0.650), significantly outperforming the baseline approach. The participant-derived framework proved particularly robust, with GPT-3.5 achieving $κ$ scores of 0.639 and 0.613 for 10-dimension and 5-dimension versions respectively. System-level evaluation showed that LLM judgements reliably identified top-performing retrieval approaches (RBO scores 0.71-0.76) while maintaining reasonable discrimination between systems (RBO 0.52-0.56). These findings suggest that LLMs can effectively evaluate educational resources when prompted with domain-specific criteria, though performance varies with framework complexity and input structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。