arXiv:2508.15250cs.CLcs.AI2025-08EMNLP被引 1

首个评估教师角色大模型伦理心理的基准,发现能力越强越易受诱导。

EMNLP: Educator-role Moral and Normative Large Language Models Profiling

  • 构建88个教师专属道德困境,实现职业化人格与伦理评估
  • 14个模型测试显示,教师型LLM抽象推理强但情绪应对弱,且能力越强越易被恶意提示操控
  • 适合教育AI安全研究者、政策制定者参考,尤其关注提示攻击风险

模拟职业(Simulating Professions, SP)使大语言模型(LLMs)能扮演专业角色,但相关心理与伦理评估仍不充分。本文提出EMNLP框架,用于教师角色的个性画像、道德发展阶段测量及软提示注入下的伦理风险评估。该框架扩展现有量表,构建88个教师专属道德困境,支持与人类教师的职业化对比。通过针对性软提示注入集,评估教师角色SP模型的合规性与脆弱性。在14个LLMs上的实验表明,教师角色模型展现更理想化、两极化的性格特征,在抽象道德推理上表现优异,但在情绪复杂情境中表现不佳;具备更强推理能力的模型对有害提示注入更敏感,揭示能力与安全之间的悖论。模型温度等超参数仅对部分风险行为有有限影响。本研究首次建立教师角色大模型伦理与心理对齐的评估基准,资源已公开:https://e-m-n-l-p.github.io/。

原文摘要 · Abstract (English)

Simulating Professions (SP) enables Large Language Models (LLMs) to emulate professional roles. However, comprehensive psychological and ethical evaluation in these contexts remains lacking. This paper introduces EMNLP, an Educator-role Moral and Normative LLMs Profiling framework for personality profiling, moral development stage measurement, and ethical risk under soft prompt injection. EMNLP extends existing scales and constructs 88 teacher-specific moral dilemmas, enabling profession-oriented comparison with human teachers. A targeted soft prompt injection set evaluates compliance and vulnerability in teacher SP. Experiments on 14 LLMs show teacher-role LLMs exhibit more idealized and polarized personalities than human teachers, excel in abstract moral reasoning, but struggle with emotionally complex situations. Models with stronger reasoning are more vulnerable to harmful prompt injection, revealing a paradox between capability and safety. The model temperature and other hyperparameters have limited influence except in some risk behaviors. This paper presents the first benchmark to assess ethical and psychological alignment of teacher-role LLMs for educational AI. Resources are available at https://e-m-n-l-p.github.io/.

教师角色伦理评估提示攻击大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。