arXiv:2606.12433cs.CYcs.CL2026-06被引 1

合成人物数据集的单一维度对齐不保证整体分布真实,需联合审计。

Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication

  • 提出独立性假设足迹(IAF)审计方法,检验属性组合的联合分布。
  • NPK数据集虽符合官方人口统计边缘分布,但三项联合分布严重失真。
  • 适合数据安全、合规审查及合成数据可信度研究者使用。

合成人物数据集常以与官方人口统计一致为信任依据,但下游用户实际使用的是年龄、性别、地区、职业、教育、姓名和机构状态等多维联合结构。仅边缘对齐无法保证联合分布真实。本文提出独立性假设足迹(IAF),基于数据卡中记录的独立属性组合进行审计:当有直接联合表时用其对比,否则通过规则推断验证。应用于一千万个韩国合成人物的NVIDIA Nemotron-Personas-Korea(NPK)数据集,发现其虽匹配韩国统计厅(KOSIS)边缘分布,但三项联合分布存在显著偏差:职业主导群体中的学历分布与韩国高等教育机构毕业生数据存在明显条件偏差;服役年龄分布与军方制度不符;女性在男性主导职业中的占比被过度平滑至均等,严格筛选结果依赖标注方式,但在直接标准化下年龄鲁棒性良好。跨六种其他地域的迁移性测试显示诊断结果具有地域依赖性,参考分类基数影响跨域标记数量。因此,合成人物数据在作为硅基样本使用前,必须结合披露锚定的联合审计。文中公开了参考清单、职业映射表、衍生指标与可复现脚本,可推广至其他合成人物资源。

原文摘要 · Abstract (English)

Synthetic persona datasets cite alignment with official demographics as a basis for trust, yet downstream users consume them as joint structures across age, sex, region, occupation, education, name, and institutional status. Marginal alignment does not imply that these joints are preserved. We propose the Independence-Assumption Footprint (IAF), an audit primitive that operates on the attribute combinations a dataset card itself documents as treated independently. For each such combination, IAF compares the synthetic joint against an external official or institutional reference, using direct joint tables where available and rule-implied checks otherwise. Applied to NVIDIA Nemotron-Personas-Korea (one million Korean synthetic personas), IAF finds that NPK aligns with KOSIS marginals while three joints fail. The major-by-occupation distribution against the KEIS graduate universe carries a large conditional mismatch. The age profile of military service is institutionally inconsistent. Female representation in male-dominated occupations is substantially over-flattened toward parity, with the strict screening verdict mapping-dependent and age-robust under direct standardisation. A transferability demonstration across six further NPK locales finds locale-dependent rather than universal diagnostics, with reference-taxonomy cardinality confounding cross-locale flag counts. For synthetic personas used as silicon samples, marginal claims must therefore be paired with disclosure-anchored joint audits before reuse. The released audit artefacts (reference manifests, occupational crosswalks, derived metrics, reproducibility scripts) instantiate this protocol on the NPK family and are released for retargeting at other synthetic persona resources.

合成数据联合审计数据可信度人物生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。