黑客可让大模型悄悄传密钥,且不易被发现。
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
- 用词汇分组法微调模型,把秘密藏在正常文本里。
- 32位密钥传输准确率达87%,多轮生成超97%。
- 输出自然流畅,人眼难辨,适合隐蔽数据泄露。
随着大语言模型(LLMs)融入敏感工作流,其泄露机密信息的风险日益增长。我们提出一种新型威胁模型TrojanStego:攻击者通过微调使LLM在不依赖输入控制的情况下,利用语言隐写术将敏感上下文信息嵌入看似自然的输出中。我们构建了一个风险因素分类体系,并据此评估该威胁的风险特征。为实现TrojanStego,我们提出一种基于词汇分区的实用编码方案,可通过微调由LLM学习。实验表明,受污染模型在保留提示上可靠地传输32位秘密,准确率达87%;通过三次生成结果多数投票,准确率超过97%。此外,模型仍保持高实用性,能规避人工检测并维持输出连贯性。这些结果揭示了一类被动、隐蔽、可行且危险的LLM数据外泄攻击新范式。
原文摘要 · Abstract (English)
As large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information. We propose TrojanStego, a novel threat model in which an adversary fine-tunes an LLM to embed sensitive context information into natural-looking outputs via linguistic steganography, without requiring explicit control over inference inputs. We introduce a taxonomy outlining risk factors for compromised LLMs, and use it to evaluate the risk profile of the threat. To implement TrojanStego, we propose a practical encoding scheme based on vocabulary partitioning learnable by LLMs via fine-tuning. Experimental results show that compromised models reliably transmit 32-bit secrets with 87% accuracy on held-out prompts, reaching over 97% accuracy using majority voting across three generations. Further, they maintain high utility, can evade human detection, and preserve coherence. These results highlight a new class of LLM data exfiltration attacks that are passive, covert, practical, and dangerous.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。