arXiv:2605.10659cs.CLcs.AI2026-05综述

数字人设能替代人类调查,但仅限于稳定属性领域。

When Can Digital Personas Reliably Approximate Human Survey Findings?

论文配图:When Can Digital Personas Reliably Approximate Human Survey Findings?
图 1 · 摘自论文原文
  • 用背景信息和历史回答构建数字人设,测试其与真实反馈的一致性。
  • 在稳定属性问题上人设表现好,主观或罕见问题则偏差大。
  • 检索增强架构提升效果,但结果更依赖人类响应结构而非模型本身。

基于大语言模型的数字人设被提议作为人类调查应答者的替代,但其可靠性边界尚不明确。本研究利用LISS面板数据,从受访者背景变量及2023年前调查历史构建人设,并与同一受访者2023年后保留答案进行对比。在四种人设架构、三种LLM及两个预测任务下,评估了人设在问题、个体、分布、公平性和聚类层面的表现。结果显示,数字人设在与稳定属性和价值观相关的领域能更好逼近人类响应分布,但在个体预测和多变量应答结构恢复上仍存局限。检索增强型架构带来最明显提升,但性能主要取决于人类响应结构:在低变异问题和常见模式下表现最佳,而在主观性强、异质性高或罕见响应上表现最差。研究为数字人设在调查研究中的适用性提供了实践指导。

原文摘要 · Abstract (English)

Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.

数字人设大模型调查研究可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。