构建10万个人物数据集,用于评估大模型的隐私与记忆能力。
ProfileFoundry: A Synthetic Person-Object Substrate for Privacy, Memory, and Tool-Use Evaluation in LLM Agent

- 基于确定性生成机制构建合成人物对象,保持跨字段一致性。
- 包含70万条事件、50万条关系边,支持长期状态追踪与验证。
- 适合研究模型对身份证据、文档关联和隐私处理能力的评估。
基础模型研究日益需要关于人的数据:用户状态、个人历史、关系、联系方式、文档及长期更新。真实用户数据难以共享、扰动、审计或负责任地再分发,而独立生成的虚假字段通常无法保持跨字段和时间的一致性以支持受控评估。我们提出ProfileFoundry,一个确定性生成器及固定参考版本,涵盖10万位成年合成人物,覆盖八个地区。每个对象包含类型化的当前快照、家庭、家族和雇主链接、对齐快照的事件、归一化的关系视图及生成溯源信息。该发布包含709,228条事件、40,338个家庭、52,491个雇主和518,564条有向关系边。我们报告了多项验证证据:选择性人口边缘比较、每对象不变性检查、发布范围内的引用与时间闭合性,以及巧合与溯源筛查。一项试点案例研究MatchDesk利用已认证的巧合、类型化事件和历史截断,评估模型能否区分可佐证的身份证据与不确定匹配。ProfileFoundry并非人口保真模型、文本渲染语料库或正式隐私机制,而是一个负责任的合成数据层,用于构建涉及记忆、隐私、文档理解、记录关联和代理状态的下游基础模型评估,同时确保每个产物背后的合成人物可追溯可审查。
原文摘要 · Abstract (English)
Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and longitudinal updates. Real user data is difficult to share, perturb, audit, or redistribute responsibly, while independently generated fake fields rarely preserve the cross-field and temporal consistency needed for controlled evaluation. We present ProfileFoundry, a deterministic generator and fixed reference release of 100,000 adult synthetic Person Objects across eight locales. Each object combines a typed current snapshot, household, family, and employer links, snapshot-aligned events, normalized relational views, and generation provenance. The release contains 709,228 events, 40,338 households, 52,491 employers, and 518,564 directed relationship edges. We report evidence in separate categories: selected population-marginal comparisons, per-object invariant checks, release-wide referential and temporal closure, and coincidence/provenance screens. A pilot case study, MatchDesk, uses certified coincidences, typed events, and history truncation to evaluate whether models distinguish corroborated identity evidence from underdetermined matches. ProfileFoundry is not a population-fidelity model, a rendered-text corpus, or a formal privacy mechanism. Instead, it is a responsible synthetic source layer for constructing downstream foundation-model evaluations involving memory, privacy, document understanding, record linkage, and agent state while keeping the synthetic person behind each artifact inspectable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。