arXiv:2512.20352cs.CLcs.AI2025-12被引 5

用双指标验证大模型主题分析可靠性,提升质性研究可信度。

Multi-LLM Thematic Analysis with Dual Reliability Metrics: Combining Cohen's Kappa and Semantic Similarity for Qualitative Research Validation

  • 结合卡帕系数与语义相似度,多轮运行评估大模型主题一致性。
  • 三款主流模型κ值均超0.8,Gemini表现最优(κ=0.907,相似度95.3%)。
  • 开源框架支持灵活配置,适合需可复现质性研究的学者使用。

质性研究面临可靠性挑战:传统人工编码需多人参与,耗时且一致性中等。本文提出一种基于多大模型的主题分析验证框架,融合集成验证与双重可靠性指标——卡帕系数(κ)用于评估编码者间一致性,余弦相似度衡量语义一致性。框架支持可配置参数(1-6个种子,温度0.0-2.0),可定制提示模板并支持变量替换,能从任意JSON格式中提取共识主题。以迷幻艺术治疗访谈转录本为案例,评估Gemini 2.5 Pro、GPT-4o与Claude 3.5 Sonnet三款模型,每模型执行六次独立运行。结果表明,Gemini表现最佳(κ=0.907,余弦相似度95.3%),其次为GPT-4o(κ=0.853,92.6%)和Claude(κ=0.842,92.1%)。三模型κ值均高于0.8,验证了多轮集成方法的有效性。框架成功提取共识主题:Gemini识别出6个主题(一致性50%-83%),GPT-4o识别5个,Claude识别4个。开源实现提供透明的可靠性度量、灵活配置及无结构依赖的共识提取,为可靠的人工智能辅助质性研究奠定方法基础。

原文摘要 · Abstract (English)

Qualitative research faces a critical reliability challenge: traditional inter-rater agreement methods require multiple human coders, are time-intensive, and often yield moderate consistency. We present a multi-perspective validation framework for LLM-based thematic analysis that combines ensemble validation with dual reliability metrics: Cohen's Kappa ($κ$) for inter-rater agreement and cosine similarity for semantic consistency. Our framework enables configurable analysis parameters (1-6 seeds, temperature 0.0-2.0), supports custom prompt structures with variable substitution, and provides consensus theme extraction across any JSON format. As proof-of-concept, we evaluate three leading LLMs (Gemini 2.5 Pro, GPT-4o, Claude 3.5 Sonnet) on a psychedelic art therapy interview transcript, conducting six independent runs per model. Results demonstrate Gemini achieves highest reliability ($κ= 0.907$, cosine=95.3%), followed by GPT-4o ($κ= 0.853$, cosine=92.6%) and Claude ($κ= 0.842$, cosine=92.1%). All three models achieve a high agreement ($κ> 0.80$), validating the multi-run ensemble approach. The framework successfully extracts consensus themes across runs, with Gemini identifying 6 consensus themes (50-83% consistency), GPT-4o identifying 5 themes, and Claude 4 themes. Our open-source implementation provides researchers with transparent reliability metrics, flexible configuration, and structure-agnostic consensus extraction, establishing methodological foundations for reliable AI-assisted qualitative research.

主题分析大模型验证质性研究可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。