arXiv:2507.22930cs.CLcs.SI2025-07中稿 · "The 17th Internat…被引 3

构建合成数据集,安全研究用户自曝隐私信息问题。

Protecting Vulnerable Voices: Synthetic Dataset Generation for Self-Disclosure Detection

  • 用大模型生成模拟的自曝隐私文本,保留真实语义特征。
  • 合成数据在可复现性、不可关联性和真实性上均表现良好。
  • 适合关注在线隐私保护的研究者和平台安全团队使用。

社交平台如 Reddit 存在大量用户自曝个人身份信息(PII)的帖子与评论。尽管这类分享可能带来积极社交互动,却也带来隐私泄露和网络危害风险。现有研究受限于缺乏开源标注数据集。为推动可复现的 PII 检测研究,本文提出一种新方法,生成可安全共享的合成自曝数据。我们构建了针对弱势群体的 19 类 PII 自曝分类体系,并基于 Llama2-7B、Llama3-8B 与 zephyr-7b-beta 三个大语言模型,通过序列化指令提示生成多文本片段的合成数据集。评估采用三项指标:一是训练模型在合成数据上的表现需与原始数据相当;二是合成数据应无法通过谷歌搜索等手段回溯至原用户;三是人类判别者无法区分合成与真实数据。相关数据与代码已公开发布于 https://netsys.surrey.ac.uk/datasets/synthetic-self-disclosure/,以支持社交媒体中 PII 隐私风险的持续研究。

原文摘要 · Abstract (English)

Social platforms such as Reddit have a network of communities of shared interests, with a prevalence of posts and comments from which one can infer users' Personal Information Identifiers (PIIs). While such self-disclosures can lead to rewarding social interactions, they pose privacy risks and the threat of online harms. Research into the identification and retrieval of such risky self-disclosures of PIIs is hampered by the lack of open-source labeled datasets. To foster reproducible research into PII-revealing text detection, we develop a novel methodology to create synthetic equivalents of PII-revealing data that can be safely shared. Our contributions include creating a taxonomy of 19 PII-revealing categories for vulnerable populations and the creation and release of a synthetic PII-labeled multi-text span dataset generated from 3 text generation Large Language Models (LLMs), Llama2-7B, Llama3-8B, and zephyr-7b-beta, with sequential instruction prompting to resemble the original Reddit posts. The utility of our methodology to generate this synthetic dataset is evaluated with three metrics: First, we require reproducibility equivalence, i.e., results from training a model on the synthetic data should be comparable to those obtained by training the same models on the original posts. Second, we require that the synthetic data be unlinkable to the original users, through common mechanisms such as Google Search. Third, we wish to ensure that the synthetic data be indistinguishable from the original, i.e., trained humans should not be able to tell them apart. We release our dataset and code at https://netsys.surrey.ac.uk/datasets/synthetic-self-disclosure/ to foster reproducible research into PII privacy risks in online social media.

隐私保护合成数据大模型自曝检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。