用投票理论检验大模型在情感分析中的独立性,发现其效果提升有限。
Examining Independence in Ensemble Sentiment Analysis: A Study on the Limits of Large Language Models Using the Condorcet Jury Theorem
- 通过多数投票机制整合不同模型预测结果
- 大模型组合仅带来微弱性能提升,表明决策缺乏独立性
- 适合关注模型协同与大模型局限的研究者
本文将康多塞陪审团定理应用于情感分析领域,检验各类大型语言模型(LLMs)与较简单自然语言处理(NLP)模型的表现。该定理指出,若个体分类器决策相互独立,多数投票机制可提升预测准确率。本研究通过在不同模型间实施多数投票机制进行实证检验,涵盖ChatGPT 4等先进大模型。结果表明,引入更大模型后性能仅获得微弱改善,暗示这些模型之间存在显著相关性,缺乏独立性。这一发现支持了如下假设:尽管复杂度高,大模型在情感分析的推理任务中并未显著优于简单模型,揭示了在高级NLP任务中模型独立性的实际局限。
原文摘要 · Abstract (English)
This paper explores the application of the Condorcet Jury theorem to the domain of sentiment analysis, specifically examining the performance of various large language models (LLMs) compared to simpler natural language processing (NLP) models. The theorem posits that a majority vote classifier should enhance predictive accuracy, provided that individual classifiers' decisions are independent. Our empirical study tests this theoretical framework by implementing a majority vote mechanism across different models, including advanced LLMs such as ChatGPT 4. Contrary to expectations, the results reveal only marginal improvements in performance when incorporating larger models, suggesting a lack of independence among them. This finding aligns with the hypothesis that despite their complexity, LLMs do not significantly outperform simpler models in reasoning tasks within sentiment analysis, showing the practical limits of model independence in the context of advanced NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。