用元认知自我进化框架,发现并修复大模型在教育、金融等领域的隐性风险。
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- 通过元认知自省识别模型潜在偏差,动态生成安全规则。
- 在14个主流模型上将越狱攻击成功率从57.8%降至显著更低水平。
- 适合关注模型安全与可控性的研究者及工业应用开发者。
保障大语言模型(LLMs)的安全性对实际部署至关重要,但现有措施常无法应对隐性、领域特定的风险。为此,我们构建了一个包含3,000个标注查询的跨领域数据集,涵盖教育、金融与管理。对14个主流LLM的评估显示,平均越狱成功率达57.8%。针对此问题,我们提出MENTOR——一种由元认知驱动的自我演化框架。MENTOR通过视角转换与后果推理等策略进行元认知自评估,识别潜在的模型错位。该框架采用单次遍历规则推理处理常规请求,并选择性触发元认知演化周期,修正残余不安全响应,将有效修正结果提炼为动态规则图,最终生成激活级控制信号用于未来推理。实验表明,MENTOR在所有测试领域均显著降低攻击成功率,优于现有安全对齐方法。代码与数据集已公开于https://anonymous.4open.science/r/MENTOR-Evo。
原文摘要 · Abstract (English)
Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000 annotated queries spanning education, finance, and management. Evaluations across 14 leading LLMs reveal a concerning vulnerability: an average jailbreak success rate of 57.8\%. In response, we propose MENTOR, a metacognition-driven self-evolution framework. MENTOR performs metacognitive self-assessment, using strategies such as perspective-taking and consequential reasoning to uncover latent model misalignments. MENTOR couples single-pass rule-guided inference for routine requests with a selectively invoked metacognitive evolution cycle that revises residual unsafe responses, distills successful corrections into a dynamic rule graph, and compiles validated rules into activation-level steering signals for future inference. Experiments demonstrate that MENTOR substantially reduces attack success rates across all tested domains and outperforms existing safety alignment methods. The code and dataset for MENTOR are available at: https://anonymous.4open.science/r/MENTOR-Evo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。