arXiv:2601.14528cs.CRcs.LG2026-01被引 2

用拓扑思想伪装提示词,发现大模型安全漏洞

LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation

  • 通过拓扑启发的提示混淆,系统性触发模型异常行为
  • 测试1.5万+提示词,1万高危案例中暴露严重安全缺陷
  • 为大模型安全防御提供可复现的分析框架,适合安全研究者

本研究提出一种受拓扑启发的提示混淆框架,以深入理解大语言模型(LLMs)的安全与可靠性漏洞。通过系统性地应用精心设计的提示词,我们展示了如何以意想不到的方式影响模型的潜在行为。实验覆盖了15,732个提示词,包括10,000个高优先级案例,测试对象涵盖LLama、Deepseek、KIMI(代码生成)及Claude。结果揭示了当前大模型防护机制的关键缺陷,强调亟需更强大的防御策略、可靠的检测方法和更强的鲁棒性。该工作提供了一个原理化的分析与缓解弱点击框架,旨在推动安全、负责任且可信的人工智能技术发展。

原文摘要 · Abstract (English)

In this study, we propose a homotopy-inspired prompt obfuscation framework to enhance understanding of security and safety vulnerabilities in Large Language Models (LLMs). By systematically applying carefully engineered prompts, we demonstrate how latent model behaviors can be influenced in unexpected ways. Our experiments encompassed 15,732 prompts, including 10,000 high-priority cases, across LLama, Deepseek, KIMI for code generation, and Claude to verify. The results reveal critical insights into current LLM safeguards, highlighting the need for more robust defense mechanisms, reliable detection strategies, and improved resilience. Importantly, this work provides a principled framework for analyzing and mitigating potential weaknesses, with the goal of advancing safe, responsible, and trustworthy AI technologies.

大模型安全提示攻击防御框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。