对比GPT-4o、Llama、DeepSeek-R1在论点分类中的表现,发现推理增强模型更优但仍有错误。
A comprehensive study of LLM-based argument classification: from LLAMA through GPT-4o to Deepseek-R1
- 采用链式思维提示,测试多款大模型在论点分类任务中的表现。
- GPT-4o在多数数据集上表现最佳,推理增强版DeepSeek-R1次之。
- 揭示了现有提示算法的局限性,适合研究论点挖掘与模型评估者参考。
论点挖掘(AM)是融合逻辑、语言学、修辞学等多学科的研究领域,旨在自动识别论点成分(如前提与主张)及其关系(支持、攻击或中立)。近年来,大语言模型(LLM)显著提升了论点语义分析效率。尽管已有多个评测基准,但公开论点分类数据集上对LLM性能的系统研究仍不足。本文对比了GPT、Llama及DeepSeek系列模型,涵盖不同版本与链式思维(Chain-of-Thoughts)增强变体,在Args.me和UKP等数据集上的表现。结果表明,ChatGPT-4o在多数基准中领先,而引入推理能力的Deepseek-R1表现优异。然而,两者仍存在错误,常见错误类型被详细分析。本研究为首个系统性分析这些数据集与提示算法的综合性工作,揭示了现有提示策略的缺陷,并指明改进方向。研究还深入剖析了可用数据集的不足,具有重要参考价值。
原文摘要 · Abstract (English)
Argument mining (AM) is an interdisciplinary research field that integrates insights from logic, philosophy, linguistics, rhetoric, law, psychology, and computer science. It involves the automatic identification and extraction of argumentative components, such as premises and claims, and the detection of relationships between them, such as support, attack, or neutrality. Recently, the field has advanced significantly, especially with the advent of large language models (LLMs), which have enhanced the efficiency of analyzing and extracting argument semantics compared to traditional methods and other deep learning models. There are many benchmarks for testing and verifying the quality of LLM, but there is still a lack of research and results on the operation of these models in publicly available argument classification databases. This paper presents a study of a selection of LLM's, using diverse datasets such as Args.me and UKP. The models tested include versions of GPT, Llama, and DeepSeek, along with reasoning-enhanced variants incorporating the Chain-of-Thoughts algorithm. The results indicate that ChatGPT-4o outperforms the others in the argument classification benchmarks. In case of models incorporated with reasoning capabilities, the Deepseek-R1 shows its superiority. However, despite their superiority, GPT-4o and Deepseek-R1 still make errors. The most common errors are discussed for all models. To our knowledge, the presented work is the first broader analysis of the mentioned datasets using LLM and prompt algorithms. The work also shows some weaknesses of known prompt algorithms in argument analysis, while indicating directions for their improvement. The added value of the work is the in-depth analysis of the available argument datasets and the demonstration of their shortcomings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。