arXiv:2412.10105cs.CL2024-12ACL被引 1

首个面向教育场景的多语言细粒度知识评测数据集,专为检验模型在大学课程中的真实掌握程度而设计。

MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset

  • 基于71本大学教材构建,覆盖8大领域、14个子领域,含3.3万概念与11.7万题干
  • 突破传统模板式评测,采用无模板、多层级(句子/段落)的精准知识探测机制
  • 适用于评估大模型在真实教学场景下的知识短板,助力安全落地教育应用

语言模型在广泛领域表现优异,但要安全有效地应用于真实教育场景,必须具备特定、细粒度的知识能力。现有填空式评测存在三大缺陷:未覆盖教育领域;聚焦低复杂度通用知识或宽泛主题,无法有效评估模型在具体学科中的知识水平;且常依赖模板,易导致模型预测偏差。本文提出MALAMUTE,首个面向教育的多语言、无模板、高度细粒度的探测数据集,由71本大学教材中的专家撰写、同行评审的题目构成,涵盖英语、西班牙语和波兰语。该数据集覆盖8个领域,每个领域最多包含14个子领域,进一步细化为概念与基于概念的题干,共包含33,361个大学课程知识点和116,887个测试题干。其精细粒度、教育导向及包含句子级与段落级题干的特点,使其成为评估语言模型课程相关知识的理想工具。我们在MALAMUTE上对掩码语言模型和因果语言模型的评估显示,尽管整体表现良好,但在具体学科中仍存在显著知识缺口,限制其在课堂中的安全使用,凸显了进一步改进的必要性。

原文摘要 · Abstract (English)

Language models (LMs) have excelled in various broad domains. However, to ensure their safe and effective integration into real-world educational settings, they must demonstrate proficiency in specific, granular areas of knowledge. Existing cloze-style benchmarks, commonly used to evaluate LMs' knowledge, have three major limitations. They: 1) do not cover the educational domain; 2) typically focus on low-complexity, generic knowledge or broad domains, which do not adequately assess the models' knowledge in specific subjects; and 3) often rely on templates that can bias model predictions. Here, we introduce MALAMUTE, a multilingual, template-free, and highly granular probing dataset comprising expert-written, peer-reviewed probes from 71 university-level textbooks across three languages (English, Spanish, and Polish). MALAMUTE is the first education-based cloze-style dataset. It covers eight domains, each with up to 14 subdomains, further broken down into concepts and concept-based prompts, totaling 33,361 university curriculum concepts and 116,887 prompts. MALAMUTE's fine granularity, educational focus, and inclusion of both sentence-level and paragraph-level prompts make it an ideal tool for evaluating LMs' course-related knowledge. Our evaluation of masked and causal LMs on MALAMUTE shows that despite overall proficiency, they have significant gaps in knowledge when examined closely on specific subjects, hindering their safe use in classrooms and underscoring the need for further development.

知识评测教育AI多语言细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。