用真实记录验证LLM写的自传,发现96.7%的场景不符事实。
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
- 用四层评分标准逐日比对自传与真实记录
- 366天中仅12天有可验证场景,96.7%未被证实
- 用本人素材生成可提升验证率至83.3%,仍有残余错误
当大语言模型被要求撰写一个人的生平时,有多少内容是真实的?我们开展了一项场景级案例审计——据我们所知,首个基于特定主体真实记录的量化审计。研究对象即本文作者:一位撰写了366天“每日一记”第一人称轶事的女性,其内容由对话式大模型生成,输入仅为模板、两天范例及每日引语,并非其原始资料。每条记录均以独立验证资料库为基础,按预设四层评分体系进行场景级审核。将验证失败率定义为未获“确认”(正向证实)的天数比例:354/366天未通过,达96.7%(威尔逊95%置信区间94.4–98.1%)。仅有12天存在可证实场景;19天(5.2%)主张内容与记录直接矛盾。主要失败模式为“根植漂移”——虚构场景中包含真实人物、雇主和地点,但该现象在不同评估者间表现不一。独立重评复现了核心结果,表明原结果未被夸大,但四分类体系可靠性仅为中等。使用当前命名模型重新生成相同内容,在相同输入下仍100%验证失败;若以主体原始资料作为生成依据,验证率显著提升至83.3%,但仍存明显偏差。本文贡献包括测量方法、可复用的审计工具(但指出弱/未验证边界不可靠),以及一种具有量化效果的增强策略。
原文摘要 · Abstract (English)
When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day "page-a-day" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。