arXiv:2506.00089cs.CYcs.AI2025-06EMNLP被引 8

用不可见干扰项让大模型输出看似合理实则错误的内容

TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents

  • 在文档中注入人眼无法察觉的伪造标记
  • 使大模型生成看似可信但实际错误的文本内容
  • 适合关注AI滥用风险与用户引导的研究者

大语言模型在推理、写作、编辑和检索方面能力快速提升,带来越来越多功能的同时,也引发了社会担忧:用户过度依赖模型,将作业、任务或敏感文档处理完全交由模型完成,缺乏实质性参与。为缓解这一问题,我们提出一种在文档中注入不可见伪标记的技术,使大模型生成看似合理实则错误的输出。基于此,我们构建了TRAPDOC框架,旨在欺骗过度依赖模型的用户。通过实证评估,我们在多个商用大模型上验证了该框架的有效性,并与多种基线方法对比。TRAPDOC为促进用户更负责任、更深入地使用语言模型提供了坚实基础。代码已开源:https://github.com/jindong22/TrapDoc。

原文摘要 · Abstract (English)

The reasoning, writing, text-editing, and retrieval capabilities of proprietary large language models (LLMs) have advanced rapidly, providing users with an ever-expanding set of functionalities. However, this growing utility has also led to a serious societal concern: the over-reliance on LLMs. In particular, users increasingly delegate tasks such as homework, assignments, or the processing of sensitive documents to LLMs without meaningful engagement. This form of over-reliance and misuse is emerging as a significant social issue. In order to mitigate these issues, we propose a method injecting imperceptible phantom tokens into documents, which causes LLMs to generate outputs that appear plausible to users but are in fact incorrect. Based on this technique, we introduce TRAPDOC, a framework designed to deceive over-reliant LLM users. Through empirical evaluation, we demonstrate the effectiveness of our framework on proprietary LLMs, comparing its impact against several baselines. TRAPDOC serves as a strong foundation for promoting more responsible and thoughtful engagement with language models. Our code is available at https://github.com/jindong22/TrapDoc.

大模型安全文本攻击用户诱导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。