arXiv:2503.04859cs.CLcs.AI2025-03被引 2

用新方法精准测量大模型编码的分析饱和度。

Codebook Reduction and Saturation: Novel observations on Inductive Thematic Saturation for Large Language Models and initial coding in Thematic Analysis

  • 提出基于DSPy框架的新型饱和度测量技术
  • 量化评估大模型初始编码的重复性与覆盖度
  • 适合做质性分析与LLM结合研究的学者

本文反思了使用大语言模型(LLMs)进行主题分析的过程,重点关注由LLMs生成的初始编码的分析饱和度问题。主题分析是一种由多个相互关联阶段构成的经典定性分析方法,其中关键步骤是初始编码,即为数据集中的离散片段打标签。饱和度是衡量定性分析有效性的重要指标,反映初始编码的重复出现程度。本文探讨了LLMs实现分析饱和的能力,并提出一种新的诱导式主题饱和度(Inductive Thematic Saturation, ITS)测量方法。该方法利用名为DSPy的编程框架,实现对ITS的精确量化评估。

原文摘要 · Abstract (English)

This paper reflects on the process of performing Thematic Analysis with Large Language Models (LLMs). Specifically, the paper deals with the problem of analytical saturation of initial codes, as produced by LLMs. Thematic Analysis is a well-established qualitative analysis method composed of interlinked phases. A key phase is the initial coding, where the analysts assign labels to discrete components of a dataset. Saturation is a way to measure the validity of a qualitative analysis and relates to the recurrence and repetition of initial codes. In the paper we reflect on how well LLMs achieve analytical saturation and propose also a novel technique to measure Inductive Thematic Saturation (ITS). This novel technique leverages a programming framework called DSPy. The proposed novel approach allows a precise measurement of ITS.

主题分析大模型饱和度定性研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。