12种模型对比揭示:生成式模型最准,但需权衡资源与可扩展性。
Comparative Insights from 12 Machine Learning Models in Extracting Economic Ideology from Political Text
- 对比12种机器学习模型在政党宣言中识别经济意识形态的表现
- GPT-4o和Gemini 1.5 Flash在细粒度与聚合层面均最优
- 零样本模型易误判,适合初探或资源受限场景
本研究系统评估了12种机器学习模型及其变体在检测经济意识形态方面的能力。以英国六次选举的政纲文本为基准,由专家与众包标注。分析涵盖生成式、微调及零样本模型在细粒度与聚合层面的表现。结果表明,GPT-4o与Gemini 1.5 Flash在所有基准上均显著优于其他模型;但存在访问门槛与资源消耗问题。微调模型表现良好,通过领域优化实现可靠替代,但依赖训练数据,难以规模化。零样本模型在识别意识形态信号时持续遇阻,常与人工标注产生负相关,说明通用知识难以胜任特定领域任务。关键发现还包括党内显著差异、微调效果随数据量提升、零样本对提示内容敏感。研究总结各模型优劣,提出政治文本自动化分析的最佳实践。
原文摘要 · Abstract (English)
This study conducts a systematic assessment of the capabilities of 12 machine learning models and model variations in detecting economic ideology. As an evaluation benchmark, I use manifesto data spanning six elections in the United Kingdom and pre-annotated by expert and crowd coders. The analysis assesses the performance of several generative, fine-tuned, and zero-shot models at the granular and aggregate levels. The results show that generative models such as GPT-4o and Gemini 1.5 Flash consistently outperform other models against all benchmarks. However, they pose issues of accessibility and resource availability. Fine-tuning yielded competitive performance and offers a reliable alternative through domain-specific optimization. But its dependency on training data severely limits scalability. Zero-shot models consistently face difficulties with identifying signals of economic ideology, often resulting in negative associations with human coding. Using general knowledge for the domain-specific task of ideology scaling proved to be unreliable. Other key findings include considerable within-party variation, fine-tuning benefiting from larger training data, and zero-shot's sensitivity to prompt content. The assessments include the strengths and limitations of each model and derive best-practices for automated analyses of political content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。