用大模型模拟研究参与者,会引发伦理与认知层面的根本问题。
'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants

- 让大模型扮演研究对象,生成看似真实的访谈数据。
- 研究人员发现模型回应缺乏真实感和上下文深度。
- 核心质疑:大模型能否真正契合质性研究的认知逻辑。
生成式模型的兴起催生了用大型语言模型(LLMs)生成合成研究数据以替代人类参与的研究范式。我们对19位质性研究者进行了访谈,了解他们对此转变的看法。起初持怀疑态度,但发现当使用相同提问方式时,模型生成的数据中出现了类似的人类叙事。然而,在多轮对话后,研究者识别出根本性局限:大模型无法体现参与者的同意与主体性,生成的回答缺乏真实触感与情境深度,并可能削弱质性研究方法的正当性。我们认为,将大模型作为研究参与者的代理,会引发‘替代效应’,其带来的伦理与认识论问题已超出当前技术限制,触及质性研究本质是否接纳此类工具的核心议题。
原文摘要 · Abstract (English)
The recent excitement around generative models has sparked a wave of proposals suggesting the replacement of human participation and labor in research and development--e.g., through surveys, experiments, and interviews--with synthetic research data generated by large language models (LLMs). We conducted interviews with 19 qualitative researchers to understand their perspectives on this paradigm shift. Initially skeptical, researchers were surprised to see similar narratives emerge in the LLM-generated data when using the interview probe. However, over several conversational turns, they went on to identify fundamental limitations, such as how LLMs foreclose participants' consent and agency, produce responses lacking in palpability and contextual depth, and risk delegitimizing qualitative research methods. We argue that the use of LLMs as proxies for participants enacts the surrogate effect, raising ethical and epistemological concerns that extend beyond the technical limitations of current models to the core of whether LLMs fit within qualitative ways of knowing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。