arXiv:2505.20910cs.CL2025-05被引 7

为大模型交互中的隐私泄露问题构建了自动化标注体系。

Automated Privacy Information Annotation in Large Language Model Interactions

  • 用大模型自动生成对话数据中的隐私片段标注
  • 覆盖24.9万条真实名查询,标注15.4万处隐私信息
  • 适合研究本地化隐私检测或开发轻量级防护工具者

用户在使用大语言模型时,若以真实身份交互,可能无意中泄露隐私信息。自动识别查询中是否存在隐私泄露及具体泄露内容已成为迫切需求。现有隐私检测方法多面向匿名内容,难以适用于真实身份交互场景。为此,我们构建了一个大规模多语种数据集,包含24.9万条用户查询和15.4万条标注的隐私片段。通过设计基于强模型的自动化标注流程,从对话数据中提取并标注隐私信息。同时,我们定义了多层次评估指标,涵盖隐私泄露、提取片段与信息类型。还建立了无需微调与基于微调的轻量级模型基线,并进行了全面性能评估。结果表明当前方法与实际应用需求仍有差距,亟需更有效的本地化隐私检测技术,本数据集为后续研究提供基础。

原文摘要 · Abstract (English)

Users interacting with large language models (LLMs) under their real identifiers often unknowingly risk disclosing private information. Automatically notifying users whether their queries leak privacy and which phrases leak what private information has therefore become a practical need. Existing privacy detection methods, however, were designed for different objectives and application domains, typically tagging personally identifiable information (PII) in anonymous content, which is insufficient in real-name interaction scenarios with LLMs. In this work, to support the development and evaluation of privacy detection models for LLM interactions that are deployable on local user devices, we construct a large-scale multilingual dataset with 249K user queries and 154K annotated privacy phrases. In particular, we build an automated privacy annotation pipeline with strong LLMs to automatically extract privacy phrases from dialogue datasets and annotate leaked information. We also design evaluation metrics at the levels of privacy leakage, extracted privacy phrase, and privacy information. We further establish baseline methods using light-weight LLMs with both tuning-free and tuning-based methods, and report a comprehensive evaluation of their performance. Evaluation results reveal a gap between current performance and the requirements of real-world LLM applications, motivating future research into more effective local privacy detection methods grounded in our dataset.

隐私保护大模型数据标注本地检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。