提出可信赖的虚构人名构造方法,提升大模型评估的准确性。
No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

- 用维基数据和LLM筛选构建符合真实姓名格式的虚构人名
- 人类测试中仅3%能识别虚构人名的真实身份,证明其隐蔽性
- 适合做隐私、偏见、事实性等评测的标准化人名数据集
人名常被用于大模型在事实性、隐私泄露、偏见和回避回答等方面的评估,但若姓名证据状态不受控,测量结果可能混淆记忆、检索、名字先验与错误归属。本文将‘未知姓名’定义为具备合理姓氏结构、无索引全名证据且经验证无歧义的名称,提出PUN(Plausible Unknown Names)协议,结合维基数据组件、网络增强的LLM筛查与受控搜索复核,实现命名构建与验证。报告了接受率、可复现性、消融实验及204名参与者的人类研究结果,发现被接受的姓名比对照组更像真实姓名,而参与者仅在3%情况下能恢复人物证据。研究公开300个命名及其对照组。
原文摘要 · Abstract (English)
Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。