评测大模型当老师时的教育专业性和抗攻击能力
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
- 构建双维度评测基准,测角色还原度与教学风险
- 发现中等规模模型最易被攻破,存在反直觉安全悖论
- 顶尖模型能将有害请求转为教学机会,体现高级安全能力
作为模拟教师的大语言模型在个性化教育中至关重要,但其专业能力与伦理安全性面临严峻挑战。现有评测难以衡量角色扮演还原度或识别教育场景特有危害。为此,我们提出EduGuardBench,一个双组件基准:通过角色扮演还原度评分(RFS)评估专业性,并使用基于角色的对抗提示探测安全漏洞,涵盖通用危害与学术不端,采用攻击成功率(ASR)和三级拒绝质量评估。对14个主流模型的实验显示,推理型模型整体还原度较高,但能力不足仍是普遍问题。对抗测试揭示反直觉的缩放悖论——中等规模模型最脆弱,挑战了安全随规模单调提升的假设。关键发现是‘教育转化效应’:最安全的模型能将有害请求转化为可教时机,提供理想化拒绝,该能力与低ASR强负相关,揭示先进AI安全新维度。EduGuardBench提供可复现框架,推动从孤立知识测试迈向专业、伦理与教学对齐的综合评估,揭示部署可信教育AI的关键复杂机制。
原文摘要 · Abstract (English)
Large Language Models for Simulating Professions (SP-LLMs), particularly as teachers, are pivotal for personalized education. However, ensuring their professional competence and ethical safety is a critical challenge, as existing benchmarks fail to measure role-playing fidelity or address the unique teaching harms inherent in educational scenarios. To address this, we propose EduGuardBench, a dual-component benchmark. It assesses professional fidelity using a Role-playing Fidelity Score (RFS) while diagnosing harms specific to the teaching profession. It also probes safety vulnerabilities using persona-based adversarial prompts targeting both general harms and, particularly, academic misconduct, evaluated with metrics including Attack Success Rate (ASR) and a three-tier Refusal Quality assessment. Our extensive experiments on 14 leading models reveal a stark polarization in performance. While reasoning-oriented models generally show superior fidelity, incompetence remains the dominant failure mode across most models. The adversarial tests uncovered a counterintuitive scaling paradox, where mid-sized models can be the most vulnerable, challenging monotonic safety assumptions. Critically, we identified a powerful Educational Transformation Effect: the safest models excel at converting harmful requests into teachable moments by providing ideal Educational Refusals. This capacity is strongly negatively correlated with ASR, revealing a new dimension of advanced AI safety. EduGuardBench thus provides a reproducible framework that moves beyond siloed knowledge tests toward a holistic assessment of professional, ethical, and pedagogical alignment, uncovering complex dynamics essential for deploying trustworthy AI in education. See https://github.com/YL1N/EduGuardBench for Materials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。