构建合成用户数据集,评测模型理解个人隐私信息的能力
PersonaBench: Evaluating AI Models on Understanding Personal Information through Accessing (Synthetic) Private User Data
- 用合成数据生成用户画像和私人文档
- 现有RAG模型在私密问题上准确率不足50%
- 适合评估个性化AI的隐私理解能力
个性化对AI助手至关重要,尤其在处理个人用户的私有AI模型场景中。关键挑战在于让模型能访问并解析用户的私有数据(如对话历史、人机交互记录、应用使用情况),以理解其生平信息、偏好及社交关系。然而,由于数据敏感性,目前缺乏公开可用的数据集来评估模型通过直接访问个人数据理解用户的能力。为此,我们提出一种合成数据生成管道,创建多样且逼真的用户档案与模拟人类活动的私人文档。基于此,我们构建PersonaBench基准,用于评估模型从模拟私有数据中理解个人信息的表现。我们采用与用户个人资料直接相关的提问,结合模型可访问的相关私有文档,测试检索增强生成(RAG)系统。结果显示,当前检索增强模型难以从用户文档中提取个人细节来回答私密问题,凸显提升个性化能力方法的迫切需求。
原文摘要 · Abstract (English)
Personalization is critical in AI assistants, particularly in the context of private AI models that work with individual users. A key scenario in this domain involves enabling AI models to access and interpret a user's private data (e.g., conversation history, user-AI interactions, app usage) to understand personal details such as biographical information, preferences, and social connections. However, due to the sensitive nature of such data, there are no publicly available datasets that allow us to assess an AI model's ability to understand users through direct access to personal information. To address this gap, we introduce a synthetic data generation pipeline that creates diverse, realistic user profiles and private documents simulating human activities. Leveraging this synthetic data, we present PersonaBench, a benchmark designed to evaluate AI models' performance in understanding personal information derived from simulated private user data. We evaluate Retrieval-Augmented Generation (RAG) pipelines using questions directly related to a user's personal information, supported by the relevant private documents provided to the models. Our results reveal that current retrieval-augmented AI models struggle to answer private questions by extracting personal information from user documents, highlighting the need for improved methodologies to enhance personalization capabilities in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。