arXiv:2506.14634cs.CLcs.AI2025-06综述被引 6

用大模型自动分类德语开放题回复,发现微调后效果才达标。

AIn't Nothing But a Survey? Using Large Language Models for Coding German Open-Ended Survey Responses on Survey Motivation

  • 对比多个大模型与提示策略,测试其在德语调查回复上的编码能力。
  • 仅微调过的模型达到可接受的预测准确率,原始模型差异大。
  • 适合需自动化处理开放题的调研者,但需权衡成本与精度。

大语言模型(LLMs)的发展为调查研究中开放题回复的自动分类提供了新可能。由于其语言理解能力,LLMs或可替代耗时的手动编码及监督学习模型的预训练。然而,现有研究多聚焦于英文、非复杂主题或单一模型,难以判断结果是否具有普适性,也未明确其分类质量与传统方法的比较。本研究以德语调查数据为例,探讨不同LLMs在编码参与调查动机类开放题回复中的表现。我们对比了多个前沿大模型和多种提示策略,并以人工专家编码作为评估基准。结果显示,各模型表现差异显著,仅微调后的模型达到满意预测性能;不同提示策略的效果取决于所用模型。此外,未微调模型在不同动机类别上分类表现不均,导致类别分布偏差。研究讨论了对编码方法研究及实质性分析的影响,并指出研究者在选择自动化分类工具时需权衡诸多因素。本研究为理解大模型在调查研究中高效、准确、可靠应用的条件提供了实证支持。

原文摘要 · Abstract (English)

The recent development and wider accessibility of LLMs have spurred discussions about how they can be used in survey research, including classifying open-ended survey responses. Due to their linguistic capacities, it is possible that LLMs are an efficient alternative to time-consuming manual coding and the pre-training of supervised machine learning models. As most existing research on this topic has focused on English-language responses relating to non-complex topics or on single LLMs, it is unclear whether its findings generalize and how the quality of these classifications compares to established methods. In this study, we investigate to what extent different LLMs can be used to code open-ended survey responses in other contexts, using German data on reasons for survey participation as an example. We compare several state-of-the-art LLMs and several prompting approaches, and evaluate the LLMs' performance by using human expert codings. Overall performance differs greatly between LLMs, and only a fine-tuned LLM achieves satisfactory levels of predictive performance. Performance differences between prompting approaches are conditional on the LLM used. Finally, LLMs' unequal classification performance across different categories of reasons for survey participation results in different categorical distributions when not using fine-tuning. We discuss the implications of these findings, both for methodological research on coding open-ended responses and for their substantive analysis, and for practitioners processing or substantively analyzing such data. Finally, we highlight the many trade-offs researchers need to consider when choosing automated methods for open-ended response classification in the age of LLMs. In doing so, our study contributes to the growing body of research about the conditions under which LLMs can be efficiently, accurately, and reliably leveraged in survey research.

大模型调查分析自动编码德语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。