用大模型生成心理模拟数据,预测成人依恋类型效果接近真实数据。
Chatting Up Attachment: Using LLMs to Predict Adult Bonds
- 用GPT-4和Claude 3模拟不同背景成人,生成依恋访谈对话。
- 仅用合成数据训练的模型性能媲美人类真实数据训练。
- 标准化后合成数据嵌入与真人更接近,适合心理研究初探者。
医学数据获取困难,限制了AI在该领域的应用。本文评估能否通过大语言模型(LLM)生成的合成数据克服这一障碍。我们使用GPT-4和Claude 3 Opus构建模拟个体,其具有不同个人背景、童年记忆和依恋风格,并参与模拟成人依恋访谈(AAI)。利用这些模拟回答训练依恋风格预测模型。评估基于9名真实人类完成相同访谈协议的转录数据集,由心理健康专家标注。结果表明,仅使用合成数据训练的模型性能可媲美人类数据训练的模型。尽管合成回答的原始嵌入空间与真实人类回答存在差异,但引入未标注的人类数据并进行简单标准化后,两者嵌入空间趋于一致。定性分析支持此调整,且标准化后的嵌入显著提升预测准确性。
原文摘要 · Abstract (English)
Obtaining data in the medical field is challenging, making the adoption of AI technology within the space slow and high-risk. We evaluate whether we can overcome this obstacle with synthetic data generated by large language models (LLMs). In particular, we use GPT-4 and Claude 3 Opus to create agents that simulate adults with varying profiles, childhood memories, and attachment styles. These agents participate in simulated Adult Attachment Interviews (AAI), and we use their responses to train models for predicting their underlying attachment styles. We evaluate our models using a transcript dataset from 9 humans who underwent the same interview protocol, analyzed and labeled by mental health professionals. Our findings indicate that training the models using only synthetic data achieves performance comparable to training the models on human data. Additionally, while the raw embeddings from synthetic answers occupy a distinct space compared to those from real human responses, the introduction of unlabeled human data and a simple standardization allows for a closer alignment of these representations. This adjustment is supported by qualitative analyses and is reflected in the enhanced predictive accuracy of the standardized embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。