用真人访谈数据评估大模型人格模拟,更真实可靠。
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
- 基于67万+真实问答对构建评估框架
- 访谈数据驱动的模型显著优于纯参数知识
- 按任务需求选方法:风格优先用检索,准确优先用时序
大语言模型模拟真实人格需以真实个人数据为依据。现有评估多依赖人口统计、问卷或短对话,缺乏对个体实际陈述的直接检验。本文提出可扩展的访谈根基评估框架,从1000位公众人物的2.3万份经验证访谈记录中提取超过67.1万条问答对,平均每人为11.5小时内容。提出包含内容相似性、事实一致性、人格契合度与知识保留率四个维度的评价体系。系统对比显示,基于真实访谈数据的方法显著优于仅依赖传记资料或模型内生知识的方法。进一步发现:检索增强方法更擅长捕捉人格风格与回答质量,而时间顺序方法在保持事实一致性和知识留存上表现更优。该框架支持按应用需求选择合适方法,为推进人格模拟研究提供实证指导。
原文摘要 · Abstract (English)
Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as proxies, but lack direct assessment against what individuals actually said. We address this gap with an interview-grounded evaluation framework for personality simulation at a large scale. We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities, each with an average of 11.5 hours of interview content. We propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention. Through systematic comparison, we demonstrate that methods grounded in real interview data substantially outperform those relying solely on biographical profiles or the model's parametric knowledge. We further reveal a trade-off in how interview data is best utilized: retrieval-augmented methods excel at capturing personality style and response quality, while chronological-based methods better preserve factual consistency and knowledge retention. Our evaluation framework enables principled method selection based on application requirements, and our empirical findings provide actionable insights for advancing personality simulation research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。