arXiv:2603.29890cs.HCcs.AI2026-03被引 1

用访谈数据生成的智能体能模拟用户对新产品概念的群体反应,但无法精准还原个体差异。

Interview-Informed Generative Agents for Product Discovery: A Validation Study

  • 基于深度访谈构建个性化智能体,模拟用户对AI新概念的评价
  • 智能体能准确反映群体响应分布,但无法复现具体个人偏好
  • 适合早期概念筛选,不适合需要个体洞察的设计场景

大型语言模型在标准化社会科学工具上表现优异,但在产品发现中的价值尚不明确。本文研究了基于访谈信息生成的智能体是否能在概念测试中模拟用户反应。通过深度工作流访谈获取知识工作者数据,构建个性化智能体,并将其对新型AI概念的评估结果与原参与者响应进行对比。结果显示,智能体虽在群体分布上校准良好,但身份识别不精确:无法复现其对应的特定个体,却能近似总体响应分布。这揭示了大模型模拟在设计研究中的潜力与局限。尽管不适合作为个体层面洞察的替代方案,但其在早期概念筛选与迭代中具有价值,只要求分布准确性即可。文章探讨了将仿真技术负责任地融入产品开发流程的启示。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong performance on standardized social science instruments, but their value for product discovery remains unclear. We investigate whether interview-informed generative agents can simulate user responses in concept testing scenarios. Using in-depth workflow interviews with knowledge workers, we created personalized agents and compared their evaluations of novel AI concepts against the same participants' responses. Our results show that agents are distribution-calibrated but identity-imprecise: they fail to replicate the specific individual they are grounded in, yet approximate population-level response distributions. These findings highlight both the potential and the limits of LLM simulation in design research. While unsuitable as a substitute for individual-level insights, simulation may provide value for early-stage concept screening and iteration, where distributional accuracy suffices. We discuss implications for integrating simulation responsibly into product development workflows.

生成式智能体产品发现用户模拟LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。