arXiv:2505.00010cs.CLcs.AI2025-05被引 1

用语言特征检测临床教育大模型的越狱行为

Jailbreak Detection in Clinical Training LLMs Using Feature-Based Predictive Models

  • 基于158次对话的2300+提示,提取4类语言特征
  • 模糊决策树模型准确率超越提示工程方法
  • 结果可解释,适合教育场景实时监控

大型语言模型(LLMs)的越狱行为威胁其在教育等敏感领域的安全使用,使用户绕过伦理约束。本研究聚焦于模拟患者交互的临床教育平台2-Sigma中的越狱检测。我们通过四种与越狱行为强相关的语言变量,对158次对话中的2300余条提示进行标注。基于提取特征训练了决策树、基于模糊逻辑的分类器、提升方法及逻辑回归等多种预测模型。结果显示,基于特征的预测模型持续优于提示工程方法,其中模糊决策树表现最佳。研究证明,基于语言特征的模型在越狱检测中既有效又可解释。建议未来探索融合提示灵活性与规则鲁棒性的混合框架,实现教育类LLM中实时、多层级的越狱监控。

原文摘要 · Abstract (English)

Jailbreaking in Large Language Models (LLMs) threatens their safe use in sensitive domains like education by allowing users to bypass ethical safeguards. This study focuses on detecting jailbreaks in 2-Sigma, a clinical education platform that simulates patient interactions using LLMs. We annotated over 2,300 prompts across 158 conversations using four linguistic variables shown to correlate strongly with jailbreak behavior. The extracted features were used to train several predictive models, including Decision Trees, Fuzzy Logic-based classifiers, Boosting methods, and Logistic Regression. Results show that feature-based predictive models consistently outperformed Prompt Engineering, with the Fuzzy Decision Tree achieving the best overall performance. Our findings demonstrate that linguistic-feature-based models are effective and explainable alternatives for jailbreak detection. We suggest future work explore hybrid frameworks that integrate prompt-based flexibility with rule-based robustness for real-time, spectrum-based jailbreak monitoring in educational LLMs.

越狱检测医疗教育可解释性特征建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。