arXiv:2507.18055cs.CLcs.CR2025-07综述被引 3

用提示词提升生成评论的多样性与隐私保护能力

Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs

  • 基于提示工程增强评论语言风格多样性
  • 发现主流LLM生成数据存在隐私泄露风险
  • 适合关注数据安全与生成质量的研究者

大型语言模型(LLMs)生成的合成数据在数据驱动应用中日益普及,既带来成本低、可扩展的优势,也引发多样性和隐私风险问题。本文针对文本类合成数据,提出一套量化评估指标,从语言表达、情感倾向和用户视角三个维度衡量多样性,从再识别风险和风格异常点两个角度评估隐私性。实验表明,当前主流LLM在生成多样化且隐私保护良好的合成数据方面存在明显短板。基于评估结果,我们设计了一种提示词驱动的方法,在保持用户隐私的前提下有效提升生成评论的语言风格多样性。

原文摘要 · Abstract (English)

The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative to real-world data to facilitate model training, its diversity and privacy risks remain underexplored. Focusing on text-based synthetic data, we propose a comprehensive set of metrics to quantitatively assess the diversity (i.e., linguistic expression, sentiment, and user perspective), and privacy (i.e., re-identification risk and stylistic outliers) of synthetic datasets generated by several state-of-the-art LLMs. Experiment results reveal significant limitations in LLMs' capabilities in generating diverse and privacy-preserving synthetic data. Guided by the evaluation results, a prompt-based approach is proposed to enhance the diversity of synthetic reviews while preserving reviewer privacy.

生成模型隐私保护多样性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。