用大模型生成按参与者数量加权的主题词云,更准确反映访谈中的核心关切。
Word Clouds as Common Voices: LLM-Assisted Visualization of Participant-Weighted Themes in Qualitative Interviews
- 利用大模型识别概念级主题,按提及的独立参与者数统计
- 在31人、155份转录稿中比传统词云更突出实际问题
- 支持自定义提示和对比分析,适合质性研究初阶快速洞察
词云常用于总结质性访谈,但传统基于频率的方法在对话场景中表现不佳:会突出填充词、忽略同义表达、割裂语义相关概念。这限制了其在早期分析阶段快速获取可解释概览的能力。本文提出ThemeClouds——一个开源可视化工具,利用大语言模型(LLM)从对话转录本生成主题加权词云。系统通过提示LLM识别语料中的概念级主题,并统计每个主题被多少名不同参与者提及,从而生成以提及广度为基础的可视化结果。研究人员可自定义提示与参数,确保透明性与控制力。基于一项比较五种录音设备配置的用户研究(31名参与者;155份转录稿,使用Whisper ASR),该方法比频率词云及主题建模基线(如LDA、BERTopic)更有效地揭示关键设备问题。论文讨论了将LLM集成至质性工作流的设计权衡,以及对可解释性与研究者自主性的意义,还提出了交互式分析如条件对比(“diff clouds”)等新可能。
原文摘要 · Abstract (English)
Word clouds are a common way to summarize qualitative interviews, yet traditional frequency-based methods often fail in conversational contexts: they surface filler words, ignore paraphrase, and fragment semantically related ideas. This limits their usefulness in early-stage analysis, when researchers need fast, interpretable overviews of what participant actually said. We introduce ThemeClouds, an open-source visualization tool that uses large language models (LLMs) to generate thematic, participant-weighted word clouds from dialogue transcripts. The system prompts an LLM to identify concept-level themes across a corpus and then counts how many unique participants mention each topic, yielding a visualization grounded in breadth of mention rather than raw term frequency. Researchers can customize prompts and visualization parameters, providing transparency and control. Using interviews from a user study comparing five recording-device configurations (31 participants; 155 transcripts, Whisper ASR), our approach surfaces more actionable device concerns than frequency clouds and topic-modeling baselines (e.g., LDA, BERTopic). We discuss design trade-offs for integrating LLM assistance into qualitative workflows, implications for interpretability and researcher agency, and opportunities for interactive analyses such as per-condition contrasts (``diff clouds'').
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。