对比三模型在长文本意见分析中的表现,发现不同场景下各有优劣。
In-Context Learning for Long-Context Sentiment Analysis on Infrastructure Project Opinions
- 用零样本和少样本方式测试大模型处理复杂基建意见的能力
- Claude 3.5 Sonnet在多变情感长文本中表现最佳,GPT-4o更稳定
- 适合关注长文本情感分析的AI应用开发者参考
大型语言模型(LLMs)在众多任务中表现优异,但在长文档处理上仍存挑战。本研究评估了三种主流LLM——GPT-4o、Claude 3.5 Sonnet和Gemini 1.5 Pro——在包含复杂、多变观点的基础设施项目长文本上的表现,涵盖零样本与少样本两种场景。结果表明:在简单短文本的零样本场景中,GPT-4o表现更佳;而在复杂、情感波动大的长文本中,Claude 3.5 Sonnet优于GPT-4o。在少样本场景下,Claude 3.5 Sonnet整体领先,而随着示例数量增加,GPT-4o展现出更强的稳定性。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved impressive results across various tasks. However, they still struggle with long-context documents. This study evaluates the performance of three leading LLMs: GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro on lengthy, complex, and opinion-varying documents concerning infrastructure projects, under both zero-shot and few-shot scenarios. Our results indicate that GPT-4o excels in zero-shot scenarios for simpler, shorter documents, while Claude 3.5 Sonnet surpasses GPT-4o in handling more complex, sentiment-fluctuating opinions. In few-shot scenarios, Claude 3.5 Sonnet outperforms overall, while GPT-4o shows greater stability as the number of demonstrations increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。