对比GPT-3.5、PaLM2和Llama2在叙事分析中的表现差异
LLM for Comparative Narrative Analysis
- 用相同提示词测试三款大模型的叙事分析能力
- 人类评估显示三模型表现存在显著差异
- 为模型选择提供多角度实证参考
本文对GPT-3.5、PaLM2和Llama2三款主流大模型进行了多视角比较叙事分析(CNA)。采用统一提示词,在相同任务下评估其输出,确保不同模型间的公平与无偏比较。研究发现,三模型对同一提示生成了显著不同的回应,表明其在理解与分析任务上的能力存在明显差异。以人类评估作为黄金标准,从四个维度分析模型表现差异,揭示了当前大模型在叙事理解层面的不一致性。
原文摘要 · Abstract (English)
In this paper, we conducted a Multi-Perspective Comparative Narrative Analysis (CNA) on three prominent LLMs: GPT-3.5, PaLM2, and Llama2. We applied identical prompts and evaluated their outputs on specific tasks, ensuring an equitable and unbiased comparison between various LLMs. Our study revealed that the three LLMs generated divergent responses to the same prompt, indicating notable discrepancies in their ability to comprehend and analyze the given task. Human evaluation was used as the gold standard, evaluating four perspectives to analyze differences in LLM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。