用大模型模拟精神科专家,分析边缘型人格障碍患者叙事,效果接近真人。
Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts
- 用人类主题分析数据训练模型,对比其生成内容与真人专家的一致性。
- 谷歌Gemini 2.5 Pro在盲评中表现与真人无异,主题重合度达58%。
- 模型能发现人忽略的主题,可减少研究偏见,适合临床辅助分析。
基于对15万字以上住院患者临床生活故事访谈的人类主导主题分析,本研究评估了大型语言模型(OpenAI的GPT、Google的Gemini、Anthropic的Claude)在支持定性临床分析中的表现。采用混合方法:研究A通过盲评和非盲评,由精神科专家从语义一致性、杰卡德系数及可信度、连贯性、结果实质性与数据根基等维度评估AI生成内容;研究B利用神经嵌入法将人类与模型生成的主题描述置于多维向量空间,客观量化两者语义与语言风格差异;研究C通过115名非专家参与,考察主题冗长程度对人类作者身份感知与内容有效性的干扰。总体结果显示,AI分析与人类解释的重合度在0%-58%之间波动,所有模型均识别出人类研究者遗漏的主题,证明其可缓解人为偏见。外部评审者无法可靠区分人类与AI生成的主题。在盲态下由高级专家评估原始数据时,Gemini 2.5 Pro的表现与人类无异,其语义嵌入也最接近人类。
原文摘要 · Abstract (English)
Building on a human-led thematic analysis of clinical life-story interviews (> 150,000 words) with inpatients with Borderline Personality Disorder, this study examines the capacity of large language models (OpenAI's GPT, Google's Gemini, and Anthropic's Claude) to support qualitative clinical analysis. The models' interpretative potential was evaluated using a mixed-methods approach. Study A involved blinded and non-blinded judges in phenomenology and clinical psychology. The experts assessed the validity of AI-generated content using semantic congruence, Jaccard coefficients, and multidimensional validity ratings, including credibility, coherence, the substantiveness of results, and grounding in qualitative data. In Study B, neural methods were used to embed human- and model-generated theme descriptions in a multidimensional vector space. This approach provided an objectified computational measure of the difference between human and model semantics and linguistic style. In Study C, complementary non-expert evaluations were conducted (N=115) to examine the influence of thematic verbosity on the perception of human authorship and content validity. Overall, the results of AI analysis showed a highly variable overlap (0-58%) with the human interpretation, while all models identified themes originally omitted by human researchers, proving their capacity to mitigate human bias. At the thematic level, external evaluators were unable to reliably distinguish human-authored themes from those generated by AI. In terms of content validity assessed against raw data by high-level experts in the blinded mode, the performance of Gemini 2.5 Pro was indistinguishable from that of humans. The comparison of semantic vector embeddings showed that its style was also the closest to humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。