大模型在文化分析中可辅助分类与理论验证,但新任务表现较弱。
On Classification with Large Language Models in Cultural Analytics
- 用提示工程对比大模型与传统方法在10个公开数据集上的分类表现。
- 大模型在已有任务上表现接近传统模型,但新任务上效果下降。
- 可作理论检验的中介输入,支持更深层的文化意义解读。
本文调研文化分析中分类作为理解实践的应用方式,评估大语言模型在此领域的适配性。我们识别出10个由公开数据集支持的任务,实证比较了大模型与传统监督方法的性能,并探索大模型在超越准确率之外的意义建构作用。结果表明,基于提示的大模型在既有任务上表现与传统模型相当,但在全新任务上表现较差。此外,大模型可通过作为正式理论检验的中介输入,辅助意义建构。
原文摘要 · Abstract (English)
In this work, we survey the way in which classification is used as a sensemaking practice in cultural analytics, and assess where large language models can fit into this landscape. We identify ten tasks supported by publicly available datasets on which we empirically assess the performance of LLMs compared to traditional supervised methods, and explore the ways in which LLMs can be employed for sensemaking goals beyond mere accuracy. We find that prompt-based LLMs are competitive with traditional supervised models for established tasks, but perform less well on de novo tasks. In addition, LLMs can assist sensemaking by acting as an intermediary input to formal theory testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。