arXiv:2411.00890cs.CLcs.AI2024-11被引 10

小模型微调后可媲美ChatGPT-4,兼顾隐私与可复现性。

Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research

  • 用小规模开源模型微调,实现高性能文本分类。
  • 微调效果随训练集增大而提升,但存在边际递减。
  • 适合注重数据隐私和研究可复现性的社科研究人员。

大型语言模型(LLMs)因其架构决定参数量与性能。社会科学家越来越多地使用LLMs进行文本分类,这类任务人工标注难以规模化。尽管闭源大模型性能更优,但存在透明度低、敏感数据泄露风险、可复现性差及依赖专有系统等问题,且成本高昂,不适用于大规模研究。相比之下,开源模型虽初始性能较弱,但可通过微调提升表现,具备本地运行(保障数据隐私)、任务定制、社区共享和可复现工作流等优势。本研究证明:经微调的小型开源模型在文本分类上可达到甚至超越ChatGPT-4的性能。我们进一步分析了训练集大小与微调效果的关系,并提出一种结合开源与闭源模型优势的混合工作流,平衡性能、透明度与可复现性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are distinguished by their architecture, which dictates their parameter size and performance capabilities. Social scientists have increasingly adopted LLMs for text classification tasks, which are difficult to scale with human coders. While very large, closed-source models often deliver superior performance, their use presents significant risks. These include lack of transparency, potential exposure of sensitive data, challenges to replicability, and dependence on proprietary systems. Additionally, their high costs make them impractical for large-scale research projects. In contrast, open-source models, although available in various sizes, may underperform compared to commercial alternatives if used without further fine-tuning. However, open-source models offer distinct advantages: they can be run locally (ensuring data privacy), fine-tuned for specific tasks, shared within the research community, and integrated into reproducible workflows. This study demonstrates that small, fine-tuned open-source LLMs can achieve equal or superior performance to models such as ChatGPT-4. We further explore the relationship between training set size and fine-tuning efficacy in open-source models. Finally, we propose a hybrid workflow that leverages the strengths of both open and closed models, offering a balanced approach to performance, transparency, and reproducibility.

大模型文本分类可复现性开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。