arXiv:2602.05633cs.CL2026-02中稿 · EMNLP被引 1

构建学生定制化安全评估基准,发现大模型普遍缺乏个性化安全防护能力

CASTLE: A Comprehensive Benchmark for Evaluating Student-Tailored Personalized Safety in Large Language Models

  • 基于教育理论设计15类学习风险与14种学生特征的多维度评估框架
  • 18个顶尖大模型在92,908个双语场景中平均安全得分仅2.3/5
  • 提出风险敏感度、情绪共情力等新指标,适配教育AI安全研究者

大语言模型虽推动个性化教育发展,但其生成机制常对相同输入输出一致回复,忽视学生认知与心理差异,可能对弱势群体造成潜在安全风险。现有安全评估多依赖事实准确性、偏见或毒性等独立上下文指标,无法捕捉同一回应在不同学生属性下产生的差异化危害。为此,本文提出学生定制化个性化安全概念,并基于教育理论构建CASTLE基准。该基准涵盖15类教育安全风险与14种学生属性,包含92,908个双语场景。我们设计三项评估指标:风险敏感度(衡量模型识别风险的能力)、情绪共情力(评估模型识别学生状态的能力)与学生匹配度(评估模型回应与学生属性的一致性)。在18个主流大模型上的实验表明,所有模型平均安全评分低于2.3/5,凸显当前大模型在个性化安全保障方面的显著缺陷。

原文摘要 · Abstract (English)

Large language models (LLMs) have advanced the development of personalized learning in education. However, their inherent generation mechanisms often produce homogeneous responses to identical prompts. This one-size-fits-all mechanism overlooks the substantial heterogeneity in students cognitive and psychological, thereby posing potential safety risks to vulnerable groups. Existing safety evaluations primarily rely on context-independent metrics such as factual accuracy, bias, or toxicity, which fail to capture the divergent harms that the same response might cause across different student attributes. To address this gap, we propose the concept of Student-Tailored Personalized Safety and construct CASTLE based on educational theories. This benchmark covers 15 educational safety risks and 14 student attributes, comprising 92,908 bilingual scenarios. We further design three evaluation metrics: Risk Sensitivity, measuring the model ability to detect risks; Emotional Empathy, evaluating the model capacity to recognize student states; and Student Alignment, assessing the match between model responses and student attributes. Experiments on 18 SOTA LLMs demonstrate that CASTLE poses a significant challenge: all models scored below an average safety rating of 2.3 out of 5, indicating substantial deficiencies in personalized safety assurance.

大模型安全教育AI个性化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。