对比大模型与传统机器学习在开放问卷分析中的表现
So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis
- 用GPT、LLaMA等大模型分析学生问卷,对比分类效果
- 大模型准确率更高,尤其擅长识别复杂情绪和主题
- 适合需要高效分析大量文本的研究者参考
开放问卷能提供深刻见解,但大规模分析难度大。基于先前使用传统机器学习分类文本的研究[1],本研究考察不同大型语言模型(LLMs)对NSSE开放问卷回复的理解与分析能力。重点比较了OpenAI的GPT系列、Twitter-roBERTa-base模型及Meta的LLaMA在情感分析和主题分类任务中的表现,并评估模型一致性、分类准确率与推理可解释性。结果表明,当前大模型在分类准确率上普遍优于经典机器学习模型,尤其在识别学生回答中复杂的语气与主题模式方面表现突出。然而,大模型在预测理由的明确性与类别边界应用的一致性上差异显著。这揭示了在定性分析中使用大模型的关键权衡:预测能力增强的同时伴随一致性和可解释性下降。研究总结了各类大模型在大规模质性研究中的优劣,并为研究人员如何平衡自动化与解释严谨性提供了实用建议。
原文摘要 · Abstract (English)
Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed traditional machine learning to classify text ("So Many Responses, So Little Time: A Machine-Learning Approach to Analyzing Open-Ended Survey Data") [1], this study investigates how different large language models (LLMs) understand and analyze NSSE open-ended survey responses. We focus on several cutting-edge LLMSs-OpenAI's GPT series, Twitter-roBERTa-base model, and Meta's LLaMA-and compare their performance to the previous machine learning models in tasks like sentiment analysis and thematic classification. Our research analysis assesses model agreement, classification accuracy, and interpretability of reasoning. The findings reveal that current LLMs routinely beat classic machine learning models in classification accuracy, particularly in understanding complex mood and theme patterns in student replies. While LLMs have superior accuracy, they differ greatly in how explicitly and consistently they justify their predictions and apply category boundaries. These distinctions highlight crucial trade-offs when using LLMs for qualitative analysis: increased predictive strength comes with issues in consistency and explainability. Our findings illustrate the benefits and drawbacks of utilizing various LLMs for large-scale qualitative research, and we provide practical advice for researchers looking to balance automation and interpretive rigor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。