用提示词提升生成评论的多样性与隐私保护能力
Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs
- 基于提示工程增强评论语言风格多样性
- 发现主流LLM生成数据存在隐私泄露风险
- 适合关注数据安全与生成质量的研究者
大型语言模型(LLMs)生成的合成数据在数据驱动应用中日益普及,既带来成本低、可扩展的优势,也引发多样性和隐私风险问题。本文针对文本类合成数据,提出一套量化评估指标,从语言表达、情感倾向和用户视角三个维度衡量多样性,从再识别风险和风格异常点两个角度评估隐私性。实验表明,当前主流LLM在生成多样化且隐私保护良好的合成数据方面存在明显短板。基于评估结果,我们设计了一种提示词驱动的方法,在保持用户隐私的前提下有效提升生成评论的语言风格多样性。
原文摘要 · Abstract (English)
The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。