发现大模型任务对话系统会泄露训练数据,可提取上千条敏感对话信息。
Extracting Training Dialogue Data from Large Language Model based Task Bots
- 设计新攻击方法,从大模型中提取任务对话训练数据。
- 最佳情况下能以超70%准确率还原数千条对话状态标签。
- 揭示隐私风险并提出针对性防护策略,适合安全与隐私研究者。
大型语言模型(LLMs)被广泛用于增强任务导向对话系统(TODS),通过建模复杂语言模式并生成上下文相关回复。然而,这种集成带来显著隐私风险:LLMs作为软知识库,可能无意中记忆包含电话号码等个人身份信息的训练对话数据,甚至完整旅行日程等对话级事件。尽管这一隐私问题至关重要,但大模型记忆如何影响任务对话机器人的发展仍未知。本文通过系统性定量研究,评估现有训练数据提取攻击,分析任务对话建模特性导致现有方法失效的原因,并提出针对基于大模型的TODS的新攻击技术,提升响应采样与成员推断能力。实验表明,所提方法在最佳情况下可提取数千条对话状态标签,精度超过70%。同时,深入分析了大模型任务对话系统中训练数据记忆的关键影响因素,并讨论了针对性缓解策略。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been widely adopted to enhance Task-Oriented Dialogue Systems (TODS) by modeling complex language patterns and delivering contextually appropriate responses. However, this integration introduces significant privacy risks, as LLMs, functioning as soft knowledge bases that compress extensive training data into rich knowledge representations, can inadvertently memorize training dialogue data containing not only identifiable information such as phone numbers but also entire dialogue-level events like complete travel schedules. Despite the critical nature of this privacy concern, how LLM memorization is inherited in developing task bots remains unexplored. In this work, we address this gap through a systematic quantitative study that involves evaluating existing training data extraction attacks, analyzing key characteristics of task-oriented dialogue modeling that render existing methods ineffective, and proposing novel attack techniques tailored for LLM-based TODS that enhance both response sampling and membership inference. Experimental results demonstrate the effectiveness of our proposed data extraction attack. Our method can extract thousands of training labels of dialogue states with best-case precision exceeding 70%. Furthermore, we provide an in-depth analysis of training data memorization in LLM-based TODS by identifying and quantifying key influencing factors and discussing targeted mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。